AI benchmark explainers

Explainers on AI benchmarks: what SWE-bench, GPQA, ARC-AGI, Humanity’s Last Exam and LMArena measure.

No stories published yet.

These pages explain what each benchmark tests, how it is run and scored, and what it misses, from the maintainers’ own documentation, so a published score can be read in context. Scores reported by labs and maintainers are collected on the LLM leaderboard, and new results are in AI benchmark news.

The Model Press

What are you looking for?

Search by headline, topic or keyword.