The models that passed Sonnet 5: Kimi K3, Grok, and Astra

The models that passed Sonnet 5: Kimi K3, Grok, and Astra
Share

The models that passed Sonnet 5: Kimi K3, Grok, and Astra

Four launches since June. None of them is a drop-in for the others. All of them make Sonnet 5 look parked.


Sonnet 5 is still a strong model. It is also a June 30 checkpoint sitting next to models that shipped on purpose after the long-horizon coding benches became the scoreboard. This is who passed it, and on what.

The predecessor to this piece is what is going wrong with Claude since Sonnet 5. Read the numbers below as vendor tables. A DeepSWE score from Moonshot's harness is not the same run as Google's model-card row.

Kimi K3: the open model that showed up in July

Moonshot released Kimi K3 on July 14 as a 2.8-trillion-parameter open model with native vision and a 1M context. It is built for long engineering sessions: large repos, terminal tools, and work that used to be reserved for closed flagships.

On DeepSWE 1.1, comparisons assembled after launch put K3 around 67.5% with Kimi Code and max thinking. Google's model card, a different harness, lists Sonnet 5 at 53.8%. Moonshot's own write-up says K3 trails only the strongest proprietary flagships. That is the right shape of the claim. An open model in July cleared the coding-agent bar Sonnet had set at the end of June, and you can run the weights.

In ThinkReview, K3 is on Professional and Teams. Use it when the review needs a long context and you want a model that is not locked to one lab's September price.

Grok 4.6, then Grok 4.7: the price-performance pass

Grok 4.6 was already the warning. Same $2 / $6 token price the 4.7 launch kept, and a coding focus aimed at long-running agents. SpaceXAI's later table scores 4.6 at 40.4% on CursorBench 4.0 and 65.2% on DeepSWE v1.1.

Grok 4.7, out September 21, is a larger base, trained longer on multi-hour tasks, at the same price. SpaceXAI's numbers: CursorBench 4.0 46.3%, DeepSWE 71.0% at high effort, GDPval 1,695, Terminal-Bench 4.0 38.0%. The CursorBench chart draws Sonnet 5 on a weaker cost curve. Fable 5.1 still leads several of those rows. Grok's point is that you no longer need Fable's budget to leave Sonnet's June scores behind.

Output tokens are $6 versus Sonnet's $10. On a review that reasons for a while, that gap is the whole invoice.

Both Groks are in ThinkReview. Grok 4.7 is the one to A/B against Sonnet on a messy multi-file diff.

GPT-6 Astra: a flagship, and still relevant

Astra is not a Sonnet competitor on price. OpenAI launched it in early September in the $10 / $50 band, next to Fable 5.1. Independent scoring from Artificial Analysis puts Astra at 61 on the Intelligence Index and 67 on the Coding Agent Index, behind Fable 5.1 at 66 and 70.

Where Astra matters for this argument is the coding-agent rows OpenAI published, including DeepSWE v1.1 around 74.1% and Terminal-Bench 4.0 at 57.7%, and the fact that SpaceXAI's GDPval chart places Astra at 1,542, behind Grok 4.7's 1,695. Astra leads the math and computer-use numbers OpenAI is proud of, including FrontierMath Tier 4. It does not uniformly beat Fable. It does sit on the 2026 scoreboard Sonnet 5 has not been updated onto.

Use Astra when the task is the flagship job: hard science, computer use, a budget that already assumed Fable. Use it as evidence that the frontier moved, not as the default reviewer.

Why these four pulled ahead of one checkpoint

They shipped into the failure mode Sonnet 5 was already aimed at, and then they shipped again.

  • Cadence. K3 in July, Grok 4.6, then 4.7 on September 21, Astra in September. Sonnet's id did not change.
  • The bench. DeepSWE, CursorBench 4.0, and Terminal-Bench 4.0 reward staying with a task. A June score on an older suite does not answer them.
  • The token. Grok matches Sonnet's $2 input and undercuts output. Gemini 3.8 Flash is cheaper still. A newer checkpoint that is also cheaper will replace a default even when the quality gap is modest.
  • Anthropic's other model. Fable 5.1 absorbed the September point release. The mid-tier inherited the June weights.

How to read a comparison without getting fooled

Rank models on one harness at a time. Google's DeepSWE row is comparable inside that card. SpaceXAI's CursorBench chart is comparable inside that chart. Moonshot's Kimi Code number is Moonshot's harness. Mixing them into one leaderboard is how a blog pretends to more precision than the labs published.

For a pull request, the honest test is narrower than any of those. Same diff, same instructions, two models. ThinkReview is built for that on GitHub, GitLab, Azure DevOps, and Bitbucket. Sonnet 5, Kimi K3, and the Groks are in Model selection.

What a Sonnet point release would have to clear is the third piece.


Sources: Kimi K3, Grok 4.7, Sonnet 5, and the Gemini 3.8 model card. Compare them on a real diff with ThinkReview.