Claude Sonnet 5.5 review results diverge sharply by configuration. CodeRabbit reports faster code reviews than Sonnet 5, while Artificial Analysis finds substantial output-token use at maximum effort. Both can be true: they measure different workloads and reasoning settings.

Research review, checked October 6, 2026. Sources are published model evaluations and an attributed practitioner experiment, not a hands-on test by The Model Press.

On this page

Anthropic’s faster everyday model

Anthropic’s September 28 announcement presents Sonnet 5.5 for well-scoped coding and everyday professional work. It reports a speed improvement of more than 30% over Sonnet 5 in its tests, while retaining Opus for more open-ended work that requires sustained judgment.

The effort setting needs to accompany any result. Anthropic lists medium as the default in its apps and Claude Code, with high on its platform. A max-effort evaluation therefore does not describe every default interaction. The model record covers release details.

A review pipeline gets faster

CodeRabbit’s published evaluation uses recorded inputs to compare review configurations. On 13 difficult Signal cases, Sonnet 5.5 with thinking on catches six known issues, against four for Sonnet 5. Actionable precision is similar: 41.2% versus 40.0%.

On a broader set of 44 pull requests, mean full-review time falls from 13 minutes 31 seconds to 6 minutes 33 seconds. Reported comments decline from 146 to 111. These are pipeline measurements, including processing beyond a single model response.

The broader run’s judge scoring was still pending in the report. Fewer comments can reduce reading work, but they cannot yet be treated as evidence of better bug coverage on that set. CodeRabbit is a commercial review vendor and model consumer; its controlled test is informative, with scope limited to its own pipeline.

Maximum effort changes the workload

Artificial Analysis reports an Intelligence Index score of 56 at max, against 58 for Opus 5.5. Its run uses roughly 193,000 output tokens per index task, around 60% more than Opus in that evaluation.

The report also notes a structured-output problem in the pre-release deployment that Anthropic fixed for the public model. Results affected by that issue should not be silently presented as a fresh public-release rerun.

A separate September 28 X experiment by Paweł Huryn records a max-effort BugHuntBench repository review taking 51 minutes and 580 turns, compared with 33 minutes and 370 turns for Sonnet 5 at max. That single run supplies no general quality verdict. It does show why a longer, more elaborate review needs a correctness check before being called an improvement.

Where the evidence points

For frequent, bounded reviews, CodeRabbit’s faster runs make Sonnet worth evaluating. For difficult correctness problems, its Signal test still gives Opus greater coverage: eight issues in Standard and ten in Max, versus Sonnet’s six.

The useful comparison preserves the repository, recorded inputs and verification process, then varies effort. Track bugs caught and missed alongside elapsed time, tokens and the comments a developer must assess. A speed gain is valuable only if the resulting review catches the failures that matter.

Evidence limits: The available tests do not establish a universal Sonnet-versus-Opus ranking. Reddit originals could not be read and are excluded; the X example is one experiment, not community consensus.