AI benchmark news
AI benchmark news: new results and evaluations of AI models, attributed to whoever ran them.

Claude Opus 5.5 tops the Intelligence Index on September 22, 2026
On September 22, 2026, Artificial Analysis scored Claude Opus 5.5 at 58 on its Intelligence Index at max effort.

ARC Prize publishes GPT-6 Astra ARC-AGI-3 results on September 3, 2026
On September 3, 2026, ARC Prize reported GPT-6 Astra at 62.7% on ARC-AGI-3 Semi-Private with its Standard harness.
Scores come from the lab’s model card or technical report, or from the benchmark maintainer’s own leaderboard, always with the benchmark version and the date. Published scores are collected on the LLM leaderboard, and what SWE-bench Verified, GPQA Diamond, ARC-AGI or Humanity’s Last Exam measure is covered in our benchmark explainers.