Opus lost the crown

Opus lost the crown
Share

Opus lost the crown

For two years, "the most powerful model" and "Opus" were the same sentence. They are not anymore.


GPT-6 Astra landed last week at $10 / $50 per million input and output tokens. Anthropic's answer at that exact price is already on the board: Claude Fable 5.1, shipped September 1. Under both of them, Claude Opus 5 is $5 / $25.

That gap is the story. Opus did not get worse. A new rung appeared above it.

Opus was the name for the top

Through the Claude 4 generation, Opus was the model you named when you meant the ceiling. Opus 4, then 4.5, 4.6, 4.7, 4.8, then Opus 5. Sonnet was the one you could afford every day. Haiku was the one you called in a loop. If a benchmark post said "we used the most powerful Claude," it meant Opus.

Sonnet 5, out June 30, was marketed as the Sonnet that could touch the previous Opus on some effort settings, at $2 / $10. That was still a story about Opus as the reference point. The flagship was the thing the mid-tier was measured against.

Fable broke the reference. Anthropic's own model table now lists Fable 5.1 above Opus 5: slower, $10 / $50, adaptive thinking, a June 2026 knowledge cutoff against Opus 5's May cutoff. The word "Opus" is still on the page. It is the middle row.

What is actually above it

Two labs now sell a model at the same startling list price, and they do not win the same rows.

Artificial Analysis scores Fable 5.1 at 66 on its Intelligence Index and 70 on its Coding Agent Index. Astra is 61 and 67. On the coding-agent index, Fable still leads the model OpenAI just shipped. On FrontierMath Tier 4 and computer-use numbers, Astra is the one OpenAI is willing to put in the launch post. SpaceXAI's Grok 4.7 table still gives Fable 5.1 the high score on CursorBench 4.0, 51.8%, with Grok 4.7 at 46.3% for a fraction of the cost per task.

So "most powerful" is no longer a single chair. Fable leads the independent aggregate and a lot of agent coding. Astra leads the math and systems tasks OpenAI chose to headline, including a reported sweep of ExploitBench. Opus leads the sentence people still say out of habit.

Grok 4.7 and Gemini 3.8 Flash complicate it further, because they are not trying to be the crown. They are trying to be the model you can run on every diff. A 73.7% DeepSWE score on a Flash card, next to Opus 5 at 74.0% in Google's own table, is how a workhorse makes a flagship look optional.

A ladder with three prices

Anthropic's public list, per million tokens, is now easy to memorize:

Model Input Output
Sonnet 5 $2 $10
Opus 5 $5 $25
Fable 5.1 $10 $50

Astra matches the top row: $10 / $50. Opus is half of that. Sonnet is a fifth of the output price.

The practical reading is that "use the most powerful model" became an expensive instruction. Pointing a review bot at Opus used to be the way to stop arguing. Pointing it at Fable or Astra doubles the Opus bill and five-times the Sonnet bill, and on some coding rows a $2 / $6 Grok is the one closer to the flagship than Sonnet is.

Opus is still the right model when you want Anthropic, you want more than Sonnet, and you do not want the Fable invoice. It is no longer the model you cite to end the argument.

What changes for review

Stop using "Opus-class" as a synonym for the ceiling. Name the rung.

  • Everyday pull requests stay on a workhorse. Sonnet 5, Grok 4.7, Kimi K3, Gemini 3.8 Flash. The crown is wasted on a three-file bugfix.
  • The hard diff can justify Opus. A cross-service auth change, a migration, a review you will only run once.
  • The new top is a separate decision. Fable or Astra, when the task is long, adversarial, or the kind of systems work those launch posts are bragging about. That choice is the subject of why this tier costs $50.

ThinkReview is built for the switch, not for one permanent crown. Pick the model in Model selection on GitHub, GitLab, Azure DevOps, and Bitbucket. Opus can still be the one you trust. It is not the one that ends the list.

The price history under this ladder — two years of GPT bills falling, then this jump — is the third piece.


Prices follow Anthropic's model overview and OpenAI's enterprise rate card for GPT-6 Astra. Coding comparisons cite Grok 4.7 and the Gemini 3.8 model card. Run the rung you mean with ThinkReview.