Claude Sonnet 5.5 is now available in ThinkReview

Claude Sonnet 5.5 is now available in ThinkReview
We're excited to announce that Claude Sonnet 5.5 is now available in ThinkReview. On September 28, Anthropic released Claude Sonnet 5.5, the second model in the Claude 5.5 family. The announcement opens with this:
Introducing Claude Sonnet 5.5, the second model in the Claude 5.5 family. It's a clear upgrade over Claude Sonnet 5, runs 30%+ faster, and costs up to 30% less for most work.
You can use it today for AI-powered code reviews on GitHub, GitLab, Azure DevOps, and Bitbucket.
What Anthropic shipped
Sonnet 5.5 is the faster, lower-cost complement to Opus 5.5. Anthropic draws the split in one sentence: Opus 5.5 is built for complex work that needs careful judgment, and Sonnet 5.5 is strongest at well-scoped everyday tasks, fixing bugs, and polished documents. Developers call it as claude-sonnet-5-5. The context window stays at 1 million tokens.
The per-token price does not move. Sonnet 5.5 is $2 per million input tokens and $10 per million output tokens, the same rates as Sonnet 5, with cache reads at $0.20. Opus 5.5 is $4 / $20. The task bill still falls, because the new model uses fewer tokens for the same work. Anthropic's testing puts that at up to 30% less per task, with outputs generating 30%+ faster. They call it the fastest Sonnet to date. In head-to-head runs, it also batched tool calls more than Sonnet 5, which meant fewer steps.
The benchmarks that matter for review
Anthropic's launch table compares Sonnet 5.5 with Sonnet 5, Opus 5.5, and GPT-6 Sol. These are the rows that map to reading a change and finishing the analysis:
| Evaluation | Sonnet 5.5 | Sonnet 5 | Opus 5.5 |
|---|---|---|---|
| Terminal-Bench 4.0 | 70.6% | 10.3% | 66.4% |
| CursorBench 4.0 | 55.5% | 34.1% | 57.8% |
| FrontierCode 1.1 (Main), Max effort | 46.2% | 42.4% | 54.4% |
| GDPval-AA v2.1 | 1844 | 1449 | 1846 |
Source: Introducing Claude Sonnet 5.5. Opus 5.5's Terminal-Bench figure is its highest score, reported at Xhigh effort. GPT-6 Sol scores 49.3% on that FrontierCode row.
Anthropic's own reading of the coding jump:
Sonnet 5.5's jump in performance is particularly noticeable in coding. At High effort on FrontierCode, it scores 10 points higher than Sonnet 5 at the same setting, at about one fifteenth of the cost per task. On CursorBench, which tests models on tasks from real Cursor coding sessions, its best score is within about two points of Opus 5.5.
Terminal-Bench 4.0 is the loud number: 70.6%, up from Sonnet 5's 10.3%, and ahead of Opus 5.5's 66.4%. The bench measures multi-step professional tasks in a command-line interface. At Medium effort, Anthropic says Sonnet 5.5 far exceeds Sonnet 5's best score for less than a tenth of the cost per task.
CursorBench 4.0 is the closer analog to a messy pull request: ambiguous, multi-file tasks taken from real coding sessions. Sonnet 5.5 scores 55.5%, against 34.1% for Sonnet 5 and 57.8% for Opus 5.5. Sualeh Asif, Director of ML at SpaceXAI, is quoted in the launch: "Claude Sonnet 5.5 delivers frontier-level performance on CursorBench 4.0 at 55.5%, second only to Opus 5.5. We think it will be a hit with developers looking to balance performance with cost."
FrontierCode asks whether an agent's code change would be merged. The Max-effort gap is narrower: 46.2% for Sonnet 5.5, 42.4% for Sonnet 5, 54.4% for Opus 5.5. Anthropic is explicit that benchmark scores are only one facet, and that in their testing and in external testing, Opus 5.5 remains clearly stronger at complex, open-ended work that needs sustained judgment.
Knowledge work moved with the code. On GDPval-AA, real tasks across 44 occupations and nine industries, Sonnet 5.5 scores 1844 — two points under Opus 5.5 at 1846, and about 400 points above Sonnet 5 at 1449.
What this means for code review
Most pull requests are the job Anthropic aimed this model at: a defined change, a way to check the result, and a comment a human reviewer will act on.
- The everyday review gets a newer Sonnet. CursorBench moving from 34.1% to 55.5% is the coding-session number. Multi-file diffs are where a June Sonnet was starting to look parked next to models that shipped through the summer.
- It can stay with a trace. Terminal-Bench going from 10.3% to 70.6% is the difference between a note on the hunk in front of you and a review that follows a failure through the steps that produced it.
- Fewer tokens, same list price. A review that reasons for a while is an output-token bill. Up to 30% less per task, generated 30%+ faster, is how a stronger Sonnet lands on more pull requests without a new rate card. Opus 5.5 stays at twice the token price.
- Review teams are already moving the middle of the queue. David Loker, VP of AI at CodeRabbit, in the announcement: "Claude Sonnet 5.5 shows better judgment than Sonnet 5 across different levels of complexity, while spending significantly fewer output tokens. Sonnet 5's tendency to reach for web search too often and its high token use are both gone in this new model. We plan to move simple and moderate reviews over now, and more in the coming weeks."
- Large changes are in scope. Daniel Vogel, chief operating officer at Epic Games, said early testing "cleared the same quality bar you'd expect from a higher-tier model, holding up on a system design audit and a data flow review," and that the model "managed tens of thousands of lines of code for gameplay system architecture."
Use Sonnet 5.5 when the question is whether this change is correct, complete, and safe to merge. Keep an Opus-class model for the review that is open-ended and has to hold a judgment across the design. Pair either one with repository-level context when the real bug lives outside the hunk.
Vendor tables are not your team's review bar. The honest comparison is still the same diff, the same instructions, and two models.
How to use Claude Sonnet 5.5 in ThinkReview
- Open ThinkReview settings — Click the extension icon and go to Settings or Model selection.
- Choose Claude Sonnet 5.5 — Select it from the model dropdown for your reviews.
- Run a review — Open any pull request or merge request and start ThinkReview with your chosen model.
Claude Sonnet 5.5 is available on Professional (and plans that include Professional models). Manage your catalog in Model Selection on the ThinkReview portal.
Where it works
Same as always: GitHub, GitLab, Azure DevOps, and Bitbucket Cloud. One extension, one workflow — now with Anthropic's Sonnet 5.5.
If you try Claude Sonnet 5.5 on real pull requests and have feedback, we'd love to hear it — open an issue on GitHub or reach out via thinkreview.dev.
Ready for faster Sonnet-class reviews? Install ThinkReview or manage your models in the portal.
Quotes, prices, and benchmark figures reference Anthropic's Claude Sonnet 5.5 announcement, September 28, 2026.