What Sonnet 5.1 would have to beat

What Sonnet 5.1 would have to beat
People want a point release because the models around Sonnet 5 did not wait.
Search for Claude Sonnet 5.1 and you will find a name. You will not find claude-sonnet-5-1 in Anthropic's model overview. The current Sonnet is still Sonnet 5, released June 30, 2026, at $2 / $10 per million tokens.
The rumor exists because the scoreboard moved and the model id did not. This piece is the rumor. The two pieces beside it are the evidence: what is going wrong since Sonnet 5, and the models that passed it.
Where the 5.1 name came from
Around August 21–24, screenshots circulated of two early-access strings: claude-marshmallow-eap and claude-melon-eap. The strings do not contain "opus", "sonnet", or "5.1". The product names were added afterward, by people reading tester impressions.
The common reading, written up by Times of AI and examined by CellCog, was Marshmallow as an Opus successor and Melon as Sonnet 5.1. A minority read Melon as a new Haiku. Testers who claimed access said neither felt like Fable. The -eap suffix is early access, a pattern Anthropic has used before a public model. That is the entire leak. No model card, no price, no date.
On September 1 Anthropic shipped Fable 5.1, and a restricted Mythos 5.1 beside it. The launch did not claim Marshmallow or Melon. No claude-sonnet-5-1 id appeared next to it. Fable taking the 5.1 slot makes a later Sonnet point release more plausible. It does not make Melon into Sonnet 5.1.
Why the rumor sticks
A point release is what the rest of the industry did to the workhorse tier all summer.
Grok went from 4.6 to 4.7 and raised CursorBench 4.0 from 40.4% to 46.3% and DeepSWE from 65.2% to 71.0% at high effort, at the same $2 / $6 price. Google's 3.8 Flash model card lists Sonnet 5 at 53.8% on DeepSWE v1.1, next to Opus 5 at 74.0% and Gemini 3.8 at 73.7%. Kimi K3 and GPT-6 Astra landed in the same window. Sonnet stayed on the June weights, with a January 2026 knowledge cutoff.
People reached for "Sonnet 5.1" because that is the smallest name that would mean Anthropic had patched the model they actually default to. Fable 5.1 patches a different budget. It is $10 / $50. It leads Artificial Analysis on the Intelligence Index (66 to Astra's 61) and on the Coding Agent Index (70 to 67). It does not update claude-sonnet-5.
What a real Sonnet 5.1 would have to clear
If Anthropic ships one, the interesting questions are already public. A blog post that invents its benchmark scores is guessing.
- DeepSWE v1.1. Sonnet 5's published figure on Google's card is 53.8%. Grok 4.7's high-effort number is 71.0%. Opus 5 is 74.0%. A point release that lands in the low 60s is a patch. One that lands with Opus is a new default.
- CursorBench 4.0 at Sonnet's price. Grok 4.7 is 46.3% at $2 / $6. Fable 5.1 is 51.8% at a much higher cost per task. Sonnet's output token is $10. A 5.1 that cannot hold the cheap part of that chart has not answered Grok.
- The id. Until
claude-sonnet-5-1is on the models page, Marshmallow and Melon are codenames. Treat release-date posts as fan fiction. - The cutoff. Sonnet 5's reliable knowledge cutoff is January 2026. Anything called 5.1 that keeps that cutoff is a behavior patch, which is still useful, and it should be described that way.
Fable 5.1 is the template for what "5.1" means inside Anthropic right now: a real launch, a model card, a price. Sonnet has not had that second event.
What to run while the name is still a guess
Do not block reviews on a codename.
- Stay on Sonnet 5 when your team has calibrated it and the diff is ordinary. It remains Anthropic's current mid-tier, and it is in ThinkReview.
- Move the hard diffs to a model that shipped after June. Grok 4.7 and Kimi K3 are the two we would put on the same merge request.
- Keep Fable and Astra in the flagship bucket. They answer each other. They do not refresh Sonnet.
- Re-test the day a Sonnet 5.1 id exists. Same rubric, same repos. A rumor is not a migration.
ThinkReview keeps that switch in Model selection, on GitHub, GitLab, Azure DevOps, and Bitbucket. The finding still waits for you to accept it.
The codenames are reported by Times of AI and CellCog. Scores are from Grok 4.7 and the Gemini 3.8 model card. Sonnet 5's docs are Anthropic's. Compare on a real PR with ThinkReview.