What happened to Mistral after Medium 3.5?

What happened to Mistral after Medium 3.5?
The lab that made open weights feel like a frontier strategy has not shipped a new general model since April.
Mistral Medium 3.5 came out on April 28, 2026. Mistral called it a frontier-class multimodal model for agents and coding, open weights under a modified MIT license, $1.50 / $7.50 per million tokens, 256k context. That was the last general model. From that day to this one is five months.
In those five months Google shipped Gemini 3.5 through 3.8 Flash, Anthropic shipped Sonnet 5 and then Fable 5.1, Moonshot shipped Kimi K3, and SpaceXAI shipped Grok 4.6 and Grok 4.7. Mistral shipped a Lean prover and a safety classifier. The general line stayed on April.
They were the promising one
Mistral's reputation was earned early. Mistral 7B in September 2023, then Mixtral, showed that a small European lab could put open weights next to models that were supposed to stay closed. For a year the release cadence was the product: a new Small, a new Large, a code model, a vision model, weights you could actually download.
December 2, 2025 was the peak of that promise. Mistral 3 put Ministral 3 and Mistral Large 3 under Apache 2.0. Large 3 was a sparse mixture-of-experts, 41 billion parameters active out of 675 billion, described as their most capable model to date. A week later Devstral 2 landed as an open code agent at 72.2% on SWE-bench Verified. We put it in ThinkReview because that was a real coding model with a real license.
March brought Mistral Small 4. April brought Medium 3.5, and Mistral's own docs now point a long list of older models at 3.5 as the replacement. Medium had become the flagship. "Medium is the new large" was already the line in 2025. By spring 2026 it was also the ceiling.
Five months is a long time on this scoreboard
A five-month-old frontier model is not a bad model. It is an old default.
The benches people are quoting now were built for the summer. Google's Gemini 3.8 model card scores 3.8 Flash at 73.7% on DeepSWE v1.1, Opus 5 at 74.0%, and Sonnet 5 — itself a June model — at 53.8%. SpaceXAI's Grok 4.7 launch quotes 71.0% at high effort and 46.3% on CursorBench 4.0, at $2 / $6. Kimi K3, a 2.8-trillion-parameter open model from July, is in that same conversation. Medium 3.5's headline number, about 77.6% on SWE-bench Verified, is a different test from a different season. Mistral has not put a new general checkpoint on the benches the others are using.
The releases that did happen are real and narrow. Leanstral 1.5 on July 2 is a formal-proof model for Lean 4. Shieldstral on August 4 is a 3B safety classifier. Both fit a lab that still publishes. Neither is a successor to Medium 3.5 for reading a pull request.
How an open lab falls behind
Mistral's advantage was never a single score. It was that the weights showed up on a schedule, under a license a company could keep, at a price under the closed flagships. Medium 3.5 still has that shape: open weights, a mid-tier price, a coding brief.
The schedule is what broke. Open weights only stay promising if the next file is newer than the closed model your competitor deployed last month. Kimi K3 is open and from July. Devstral 2 is open and from December 2025. A team that wanted a European open model for agents in September is being pointed at an April checkpoint while Grok and Gemini Flash revised themselves three and four times.
There is a second, quieter slide. Mistral's catalog is busy retiring names into Medium 3.5. That is sensible product hygiene when 3.5 is the best model you have. It also advertises that 3.5 is not a stepping stone. It is the destination they are willing to publish.
None of this says the company stopped. Le Chat, Vibe, and the API are still there. It says the thing developers used Mistral for — a current open model you could bet a workflow on — has gone stale in a season when everyone else treated the workhorse as something you patch.
What to use while that file is missing
Keep Devstral 2 if your review stack is built on it and the diffs are the kind SWE-bench was measuring. It is still in ThinkReview on Professional and Teams.
For the long-horizon benches this summer actually moved, the models that shipped are elsewhere. Kimi K3 is the open one from July. Grok 4.7 is the cheap closed one from September. Gemini 3.8 Flash is the workhorse Google kept revising. Sonnet 5 is the June default we already wrote about, and it has the same problem in miniature: a strong launch, then a quiet model id, while the field kept going.
If Mistral ships a general model after Medium 3.5, the test is simple. Same diff, same instructions, next to Kimi and Grok. Until that model has a name and a date after April, the promising open lab is a catalog entry, not a default.
ThinkReview still lets you pick. Model selection on GitHub, GitLab, Azure DevOps, and Bitbucket. Use the April model when you mean to. Don't use it because it used to be the one that was about to catch up.
Dates and product names follow Mistral Medium 3.5, Mistral 3, Mistral Small 4, Leanstral 1.5, and Shieldstral. Compare a current model on a real PR with ThinkReview.