Artificial Analysis said on September 22, 2026 that Claude Opus 5.5 had taken the top spot on its Intelligence Index, with a score of 58 at max effort. In its Claude Opus 5.5 analysis, Artificial Analysis calls that “the highest score we have measured by several points.”
On this page
The headline score
Artificial Analysis reports that “at max effort it scores 58 on the Artificial Analysis Intelligence Index.” It also writes that Claude Opus 5.5 “brings Anthropic to parity with GPT-6 Astra on evaluations.”
The model has five effort settings (low, medium, high, xhigh and max). According to Artificial Analysis, four of the five “sit on the Intelligence vs Cost per Task frontier.”
Results on individual evaluations
Artificial Analysis published several component results alongside the index score. All figures are from its own testing.
| Evaluation | Claude Opus 5.5 | Previous reference given by Artificial Analysis |
|---|---|---|
| AA-Briefcase | 1,822 Elo | — |
| GDPval-AA v2.1 | 1,846 Elo | — |
| SciCode | 66.9% | 63.1% (Claude Fable 5.1) |
| Humanity’s Last Exam | 61.4% | 59.1%, previous best (Claude Fable 5.1) |
| Terminal-Bench 4.0 | 59.6% | level with GPT-6 Astra |
AA-Briefcase is described by the firm as “our private frontier knowledge work evaluation,” on which the model “reaches an Elo of 1822.” On Terminal-Bench 4.0, Artificial Analysis says the 59.6% score is “level with the leader GPT-6 Astra.”
How to read the numbers
The score of 58 is an Artificial Analysis Intelligence Index result, so it is comparable with other models the firm has measured on the same index, not with scores a lab publishes for its own models. The component scores above are likewise Artificial Analysis’s runs, not figures from Anthropic.
The comparisons in the table are the ones Artificial Analysis itself draws. Two of them point to Claude Fable 5.1, which held the previous best on Humanity’s Last Exam in the firm’s measurements; the Terminal-Bench 4.0 comparison is with GPT-6 Astra.
What it changes
By Artificial Analysis’s account, the September 22 result puts an Anthropic model at the top of its index, by what the firm describes as several points. On the evaluations it highlights, Claude Opus 5.5 sets a new best on Humanity’s Last Exam in the firm’s measurements and matches the leader on Terminal-Bench 4.0.
Claude Opus 5.5’s scores are recorded with their source and date on our benchmark leaderboard, and the model’s release is in the AI model release timeline.
Sources
- Claude Opus 5.5 takes the top spot on the Artificial Analysis Intelligence Index — Artificial Analysis




