LLM leaderboard: benchmark scores from labs and maintainers
An LLM leaderboard of benchmark scores exactly as labs, maintainers and evaluators published them.
| Name | Model | Benchmark | Score | Unit | Setting | Reported by | Publisher | Published | Source | Lab | Checked |
|---|---|---|---|---|---|---|---|---|---|---|---|
| GPT-6 Astra (high) — ARC-AGI-3 (ARC Prize, Provider Adapter harness) | GPT-6 Astra (high) | ARC-AGI-3 | 99.9 | % | High effort; Provider Adapter harness; Semi-Private set | Benchmark maintainer | ARC Prize Foundation | arcprize.org | OpenAI | ||
| GPT-6 Astra (max) — ARC-AGI-3 (ARC Prize, Standard harness) | GPT-6 Astra (max) | ARC-AGI-3 | 62.7 | % | Max effort; Standard harness; Semi-Private set | Benchmark maintainer | ARC Prize Foundation | arcprize.org | OpenAI | ||
| Claude Opus 5.5 (max) — Artificial Analysis Intelligence Index (Sept 22 article) | Claude Opus 5.5 (max) | Artificial Analysis Intelligence Index | 58 | Index points | Max effort; index as of AA article dated 2026-09-22 | Independent evaluator | Artificial Analysis | artificialanalysis.ai | Anthropic | ||
| Gemini 3.8 Flash (high) — Artificial Analysis Intelligence Index (Sept 30 article) | Gemini 3.8 Flash (high) | Artificial Analysis Intelligence Index | 41 | Index points | High reasoning; index as of AA article dated 2026-09-30 | Independent evaluator | Artificial Analysis | artificialanalysis.ai | Google DeepMind | ||
| Gemini 4 Argon (high) — Artificial Analysis Intelligence Index (Sept 30 article) | Gemini 4 Argon (high) | Artificial Analysis Intelligence Index | 53 | Index points | High reasoning; index as of AA article dated 2026-09-30 | Independent evaluator | Artificial Analysis | artificialanalysis.ai | Google DeepMind | ||
| GPT-6 Astra (max) — Artificial Analysis Intelligence Index (Sept 30 article) | GPT-6 Astra (max) | Artificial Analysis Intelligence Index | 53 | Index points | Max effort; index as of AA article dated 2026-09-30 | Independent evaluator | Artificial Analysis | artificialanalysis.ai | OpenAI | ||
| GPT-6.1 Sol (max) — Artificial Analysis Intelligence Index (Sept 30 article) | GPT-6.1 Sol (max) | Artificial Analysis Intelligence Index | 52 | Index points | Max effort; index as of AA article dated 2026-09-30 | Independent evaluator | Artificial Analysis | artificialanalysis.ai | OpenAI | ||
| Claude Opus 5.5 — AutomationBench (Anthropic, self-reported) | Claude Opus 5.5 | AutomationBench | 40 | % | As reported in Anthropic's table | Lab (self-reported) | Anthropic | anthropic.com | Anthropic | ||
| Gemini 4 Argon — AutomationBench (Google, self-reported) | Gemini 4 Argon | AutomationBench | 51.3 | % | As reported in Google's announcement | Lab (self-reported) | blog.google | Google DeepMind | |||
| Claude Sonnet 5.5 (max) — AutomationBench-AA (Artificial Analysis) | Claude Sonnet 5.5 (max) | AutomationBench-AA | 71 | % | Max effort (Artificial Analysis run) | Independent evaluator | Artificial Analysis | artificialanalysis.ai | Anthropic | ||
| Gemini 4 Argon — AutomationBench-AA (Artificial Analysis) | Gemini 4 Argon | AutomationBench-AA | 78 | % | Artificial Analysis run | Independent evaluator | Artificial Analysis | artificialanalysis.ai | Google DeepMind | ||
| GPT-5.6 Sol — BrowseComp (OpenAI, self-reported) | GPT-5.6 Sol | BrowseComp | 90.4 | % | As reported in OpenAI's announcement | Lab (self-reported) | OpenAI | openai.com | OpenAI | ||
| GPT-6 Astra — BrowseComp (OpenAI, self-reported) | GPT-6 Astra | BrowseComp | 91.5 | % | Maximum across effort levels | Lab (self-reported) | OpenAI | openai.com | OpenAI | ||
| Kimi K3 (max) — BrowseComp (Moonshot AI, 1M-token context) | Kimi K3 (max) | BrowseComp | 90.4 | % | Max reasoning effort; 1M-token context, no context management | Lab (self-reported) | Moonshot AI | kimi.ai | Moonshot AI | ||
| Claude Opus 5.5 — CursorBench 4.0 (Anthropic, self-reported) | Claude Opus 5.5 | CursorBench 4.0 | 57.8 | % | As reported in Anthropic's table | Lab (self-reported) | Anthropic | anthropic.com | Anthropic | ||
| Claude Sonnet 5.5 — CursorBench 4.0 (Anthropic, self-reported) | Claude Sonnet 5.5 | CursorBench 4.0 | 55.5 | % | As reported in Anthropic's table | Lab (self-reported) | Anthropic | anthropic.com | Anthropic | ||
| Grok 4.7 — CursorBench 4.0 (xAI, self-reported) | Grok 4.7 | CursorBench 4.0 | 46.3 | % | As reported in xAI's announcement | Lab (self-reported) | xAI | x.ai | xAI | ||
| Gemini 4 Argon — DeepSWE v1.1 (Google, self-reported) | Gemini 4 Argon | DeepSWE v1.1 | 77.9 | % | As reported in Google's announcement | Lab (self-reported) | blog.google | Google DeepMind | |||
| GPT-6 Luna (max) — DeepSWE v1.1 (OpenAI, self-reported) | GPT-6 Luna (max) | DeepSWE v1.1 | 66.6 | % | Max effort | Lab (self-reported) | OpenAI | openai.com | OpenAI | ||
| GPT-6 Sol (max) — DeepSWE v1.1 (OpenAI, self-reported) | GPT-6 Sol (max) | DeepSWE v1.1 | 68.8 | % | Max effort | Lab (self-reported) | OpenAI | openai.com | OpenAI | ||
| Grok 4.7 (high) — DeepSWE v1.1 (xAI, self-reported) | Grok 4.7 (high) | DeepSWE v1.1 | 71 | % | High effort | Lab (self-reported) | xAI | x.ai | xAI | ||
| Kimi K3 (max) — DeepSWE v1.1 (Moonshot AI, mini-SWE-agent harness) | Kimi K3 (max) | DeepSWE v1.1 | 67.3 | % | Max reasoning effort; mini-SWE-agent harness; temperature 1.0, top-p 1.0 | Lab (self-reported) | Moonshot AI | kimi.ai | Moonshot AI | ||
| GPT-6 Astra — FrontierMath Tier 4 (OpenAI, self-reported) | GPT-6 Astra | FrontierMath Tier 4 | 97.6 | % | FrontierMath Tier 4 v2; maximum across effort levels | Lab (self-reported) | OpenAI | openai.com | OpenAI | ||
| Claude Opus 5.5 (max) — GDPval-AA v2.1 (Artificial Analysis) | Claude Opus 5.5 (max) | GDPval-AA v2.1 | 1,846 | Elo | Max effort (Artificial Analysis run) | Independent evaluator | Artificial Analysis | artificialanalysis.ai | Anthropic | ||
| GPT-6 Astra — GPQA Diamond (OpenAI, self-reported) | GPT-6 Astra | GPQA Diamond | 96 | % | Maximum across effort levels (research environment) | Lab (self-reported) | OpenAI | openai.com | OpenAI | ||
| Kimi K3 (max) — GPQA Diamond (Moonshot AI model card) | Kimi K3 (max) | GPQA Diamond | 93.5 | % | Max reasoning effort; temperature 1.0 | Lab (self-reported) | Moonshot AI | huggingface.co | Moonshot AI | ||
| Gemini 3.8 Flash — HLE-Verified (Google, self-reported) | Gemini 3.8 Flash | HLE-Verified | 54.9 | % | As reported in Google's announcement | Lab (self-reported) | blog.google | Google DeepMind | |||
| Claude Fable 5.1 — Humanity's Last Exam, no tools (Anthropic, self-reported) | Claude Fable 5.1 | Humanity's Last Exam | 60.9 | % | No tools | Lab (self-reported) | Anthropic | anthropic.com | Anthropic | ||
| Claude Opus 5.5 — Humanity's Last Exam, with tools (Anthropic, self-reported) | Claude Opus 5.5 | Humanity's Last Exam | 67.7 | % | With tools | Lab (self-reported) | Anthropic | anthropic.com | Anthropic | ||
| Claude Opus 5.5 (max) — Humanity's Last Exam (Artificial Analysis) | Claude Opus 5.5 (max) | Humanity's Last Exam | 61.4 | % | Max effort (Artificial Analysis run) | Independent evaluator | Artificial Analysis | artificialanalysis.ai | Anthropic | ||
| Gemini 4 Argon — LVBench (Google, self-reported) | Gemini 4 Argon | LVBench | 91.7 | % | As reported in Google's announcement | Lab (self-reported) | blog.google | Google DeepMind | |||
| GPT-6 Astra — OSWorld 2.0 (OpenAI, self-reported) | GPT-6 Astra | OSWorld 2.0 | 72.6 | % | Maximum across effort levels | Lab (self-reported) | OpenAI | openai.com | OpenAI | ||
| Claude Opus 5.5 — OSWorld 2.1 (Anthropic, self-reported, partial) | Claude Opus 5.5 | OSWorld 2.1 | 81.8 | % | Partial-credit scoring | Lab (self-reported) | Anthropic | anthropic.com | Anthropic | ||
| Claude Sonnet 5.5 — OSWorld 2.1 (Anthropic, self-reported, partial) | Claude Sonnet 5.5 | OSWorld 2.1 | 80.1 | % | Partial-credit scoring | Lab (self-reported) | Anthropic | anthropic.com | Anthropic | ||
| Claude Opus 5.5 (max) — SciCode (Artificial Analysis) | Claude Opus 5.5 (max) | SciCode | 66.9 | % | Max effort (Artificial Analysis run) | Independent evaluator | Artificial Analysis | artificialanalysis.ai | Anthropic | ||
| GPT-5.6 Sol — Terminal-Bench 2.1 (OpenAI, self-reported) | GPT-5.6 Sol | Terminal-Bench 2.1 | 88.8 | % | As reported in OpenAI's announcement | Lab (self-reported) | OpenAI | openai.com | OpenAI | ||
| Kimi K3 (max) — Terminal-Bench 2.1 (Moonshot AI model card) | Kimi K3 (max) | Terminal-Bench 2.1 | 88.3 | % | Max reasoning effort; Kimi Code harness | Lab (self-reported) | Moonshot AI | huggingface.co | Moonshot AI | ||
| Claude Fable 5.1 — Terminal-Bench 4.0 (Anthropic, self-reported) | Claude Fable 5.1 | Terminal-Bench 4.0 | 55.8 | % | As reported in Anthropic's table | Lab (self-reported) | Anthropic | anthropic.com | Anthropic | ||
| Claude Opus 5.5 (max) — Terminal-Bench 4.0 (Artificial Analysis) | Claude Opus 5.5 (max) | Terminal-Bench 4.0 | 59.6 | % | Max effort (Artificial Analysis run) | Independent evaluator | Artificial Analysis | artificialanalysis.ai | Anthropic | ||
| Claude Opus 5.5 (xhigh) — Terminal-Bench 4.0 (Anthropic, self-reported) | Claude Opus 5.5 (xhigh) | Terminal-Bench 4.0 | 66.4 | % | xhigh effort | Lab (self-reported) | Anthropic | anthropic.com | Anthropic | ||
| Claude Sonnet 5.5 — Terminal-Bench 4.0 (Anthropic, self-reported) | Claude Sonnet 5.5 | Terminal-Bench 4.0 | 70.6 | % | As reported in Anthropic's announcement | Lab (self-reported) | Anthropic | anthropic.com | Anthropic | ||
| Claude Sonnet 5.5 (max) — Terminal-Bench 4.0 (Artificial Analysis) | Claude Sonnet 5.5 (max) | Terminal-Bench 4.0 | 64 | % | Max effort (Artificial Analysis run) | Independent evaluator | Artificial Analysis | artificialanalysis.ai | Anthropic | ||
| Gemini 4 Argon — Terminal-Bench 4.0 (Artificial Analysis) | Gemini 4 Argon | Terminal-Bench 4.0 | 57 | % | Artificial Analysis run | Independent evaluator | Artificial Analysis | artificialanalysis.ai | Google DeepMind | ||
| GPT-6 Astra — Terminal-Bench 4.0 (OpenAI, self-reported) | GPT-6 Astra | Terminal-Bench 4.0 | 57.9 | % | Maximum across effort levels | Lab (self-reported) | OpenAI | openai.com | OpenAI | ||
| Grok 4.7 — Terminal-Bench 4.0 (xAI, self-reported) | Grok 4.7 | Terminal-Bench 4.0 | 37.6 | % | As reported in xAI's announcement | Lab (self-reported) | xAI | x.ai | xAI |
This LLM leaderboard lists benchmark scores for current AI models exactly as they were published: by the lab in its model announcement or model card, by the benchmark’s maintainer (such as the ARC Prize Foundation), or by an independent evaluator (such as Artificial Analysis). Each row is one score for one model on one benchmark, with the setting the source states, the publisher, the publication date and a link to the source page.
Read the rows with four caveats:
- Self-reported and maintainer-run scores are not directly comparable. A lab picks its own harness, prompts and effort level; a maintainer or independent evaluator runs every model the same way. The “Reported by” column tells you which is which.
- Settings differ. Effort level, tools, harness and context limits change results, sometimes by many points. The “Setting” column copies what the source says.
- Versions matter. Terminal-Bench 2.1 and 4.0, or OSWorld 2.0 and 2.1, are different tests. An index such as the Artificial Analysis Intelligence Index can also change between articles, so compare index values only within the same article date.
- Every row links its source. If a number is not on the publisher’s own page, it is not here.
The table is sorted by benchmark; filter by benchmark, lab or source type, or sort by score within one benchmark. Rows are re-checked against their sources three times a week. Model release dates are in AI models, new results are covered in AI benchmark news, and how we pick sources is in our editorial standards.
Source
Every score is copied from the lab’s announcement or model card, the benchmark maintainer’s page or the independent evaluator’s article, linked on each row with the date it was checked. How we pick and check sources: editorial standards.