AI benchmark news

AI benchmark news: new results and evaluations of AI models, attributed to whoever ran them.

Scores come from the lab’s model card or technical report, or from the benchmark maintainer’s own leaderboard, always with the benchmark version and the date. Published scores are collected on the LLM leaderboard, and what SWE-bench Verified, GPQA Diamond, ARC-AGI or Humanity’s Last Exam measure is covered in our benchmark explainers.

The Model Press

What are you looking for?

Search by headline, topic or keyword.