Self-Building Benchmarks
LLMs auto-generate and grade their own work-capability exams across occupations; even top models score only 65-79%, though performance jumped 26 points from 2024 to 2025.
Work in progress. LLMs are used to automatically generate and evaluate practical work exams for tasks across Finance & Business Operations, Management, and Computer & Mathematics occupations, extending the “LLM-as-a-judge” paradigm to occupational task assessment. Restricting to text-only tasks needing no external tools, only 7% of tasks in these occupations turned out to be testable this way. Across 13 models tested, leading models score a median of 65-79% even on these basic tasks, struggling most with data manipulation and financial calculations — but models released in 2025 averaged 66%, up from 40.5% in 2024. Validation and extending to tool-using tasks are ongoing. Read more
