The $50 token is a security product

The $50 token is a security product
Share

The $50 token is a security product

The new flagships are not a nicer Opus. They are priced like a specialist, and the labs are fencing the dangerous part.


Opus lost the top of the ladder last week, when GPT-6 Astra showed up at the same list price Anthropic already charges for Fable 5.1: $10 per million input tokens and $50 per million output tokens. That is the public price of the new ceiling. The more interesting price is the one next to it, the one you cannot click in a normal console.

OpenAI's enterprise rate card lists Daybreak Red at $12.50 / $75, and a gpt-5.6-cyber model at the same $12.50 / $75. Daybreak access is an approval program, Trusted Access for Cyber, not a dropdown. Anthropic ships Mythos 5.1 beside Fable 5.1 and marks it limited availability. Google's September Flash launch split the same way: 3.8 Flash for everyone, 3.8 Flash Cyber only through the Fairwind program.

The pattern is no longer "one model, turn the safety down." It is a second product, thinner access, higher price, aimed at security work.

What the labs are actually selling

Astra's launch posts lead with saturated math, computer use, and a reported 100% on ExploitBench. That last number is a vendor claim about a security benchmark, not a permission slip. Fable 5.1 leads the independent Intelligence Index (66 to Astra's 61) and the Coding Agent Index (70 to 67). Mythos, in the same family, is the model Anthropic does not hand to every API key.

Read those together and the $50 output token makes sense. A general flagship that is strong at long agent runs will also be strong at the tasks security teams and attackers both want: reading a codebase, finding a weak boundary, proposing a patch. The labs' answer is to sell the capability twice. Once, expensively, in public. Again, more expensively or not at all, inside a program with a name.

Gemini already rehearsed the split at the cheap end. Flash Cyber is not a pricier 3.8 for chat. It is a defender SKU with an application form. OpenAI and Anthropic are doing it at the top of the bill.

Why this tier showed up now

The workhorse got too good to be the place you hide the dangerous skills.

Grok 4.7 is $2 / $6 and posts 71.0% on DeepSWE at high effort. Gemini 3.8 Flash is $0.75 / $3.75 through December and 73.7% on the same family of long-horizon coding tests in Google's card. Sonnet 5 is $2 / $10. If those models are what you use to review a payments diff, the flagship has to justify itself with work those models are fenced off from, or with depth you only need a few times a month.

Security is that justification. A cyber program can charge $75 per million output tokens because the buyer is a team that already pays for audits, and because the lab can refuse the other buyers. The public $50 model is the version of that product that still has a credit card form.

This is the opposite of the last two years of GPT pricing, which is a separate story. The default model got cheaper. The new model got expensive on purpose.

What it means in a year

Expect the catalog to keep splitting.

  • A cheap public line that revises every few weeks. Flash, Grok, Sonnet, Kimi. This is where pull request review lives.
  • A named flagship at roughly $10 / $50. Fable and Astra are the first pair to occupy that cell at the same time. The next lab will want the cell, not a price under it.
  • A gated cyber SKU above or beside the flagship. Daybreak Red, Mythos, Flash Cyber. Access is the product. The token price is how you know you are not in the public tier.

The failure mode is pointing the gated model at every merge request. You will pay flagship prices for comments a $6 model would have written, and you will wait on an approval program for a CSS change. The other failure mode is pretending the cheap model is cleared for the work the program exists to hold. A dependency bump that touches auth, a sandbox escape, a signing-key path: that is the diff where the expensive model earns the invoice. A linter nit is not.

How to spend it on review

Keep the default on a workhorse you can run all day. In ThinkReview that is Grok 4.7, Kimi K3, Gemini 3.8 Flash, or Sonnet 5, chosen in Model selection.

Escalate when the change is security-shaped: authentication, cryptography, isolation, supply-chain scripts, anything you would hand a human security reviewer. Then the $50 model is a second opinion, not the house style. You still accept or dismiss the finding. ThinkReview does not post it until you do, on GitHub, GitLab, Azure DevOps, and Bitbucket.

The crown moved. The bill moved with it. The review workflow should move only the diffs that are worth the new price.


Public prices: Fable 5.1 and Opus 5 from Anthropic; Astra, Daybreak Red, and gpt-5.6-cyber from OpenAI's enterprise rate card. Cyber programs are access products, not review instructions. Compare models on a real diff with ThinkReview.