Seshat Global History Benchmark
A 36,000-question benchmark testing how well LLMs understand 600+ historical societies from the Seshat Databank.
Global History Benchmark. We build a benchmark using the Seshat Databank—an interdisciplinary effort , where over a decade of work went into interpreting books and articles to compile facts about 600+ historical societies. We tested LLMs on 36,000+ questions derived from the databank. In a four-choice format, LLMs achieved balanced accuracy between 33.6% (LLama-3.1-8B) and 46% (GPT-4-Turbo)—better than random guessing (25%) but far from expert comprehension. LLMs perform better on earlier historical periods. Regionally, performance is better for the Americas and lowers in Oceania and Sub-Saharan Africa for the more advanced models. Read more
Seshat-derived history questions also contributed to Humanity’s Last Exam, published in Nature.
