Composer had the editor. Grok took the scores.

Composer had the editor. Grok took the scores.
Share

Composer had the editor. Grok took the scores.

Cursor's own model had a six-month run. The benches after the SpaceX deal belong to Grok.


Cursor did not become interesting because it trained a model. It became interesting because the editor was already where the work happened. Composer was the second act: Cursor's first agentic coding model, shipped in late 2025, then climbed a very short ladder.

Composer 1.5 scaled reinforcement learning by more than 20×. Composer 2 added continued pretraining and, on Cursor's account, reached frontier-level coding at a fraction of the usual cost. Composer 2.5 followed in May. By July, the model people on Cursor's team were treating as a step past Composer was Grok 4.5.

Composer 2.5 is still in Cursor's model pool, beside Grok 4.5 and Grok 4.6. The menu kept it. The frontier left.

The partnership was a training deal

In April, Cursor and SpaceXAI agreed to train together on Colossus. Cursor's description of that moment is about a bottleneck: every larger training run had made Composer more capable, and they had run out of compute they could buy on their own.

On June 16, SpaceX agreed to acquire Anysphere, the company behind Cursor, for $60 billion in stock. The deal closed on August 14. Cursor's post that day points at Grok 4.6, released the previous Wednesday, as the first look at what the combined stack can ship.

SpaceX was buying distribution and a training loop. Coding prompts, iteration, and architecture decisions from the editor were the point of the filing language around Grok. The IDE is where the model gets used, and where the next model gets its examples.

Grok 4.5 is the handoff

Grok 4.5 is the first named result of that training relationship. SpaceXAI's launch post says it was trained alongside Cursor, aimed at coding, agents, and knowledge work, and priced at $2 / $6 per million input and output tokens. Michael Truell called it a significant step up over every model Cursor had built, Composer 2.5 included, and a daily driver for many people on the Cursor team.

The charts on that post are the public picture of the handoff. They are competitive, and they are specific.

On DeepSWE 1.0, run by Artificial Analysis with each lab's own harness, Grok 4.5 scores 62.0%. Fable max is at 66.1%. GPT 5.5 xhigh is at 64.31%. Opus 4.8 max is at 55.75%.

DeepSWE 1.0 scores from SpaceXAI's Grok 4.5 announcement, with Grok 4.5 at 62%

Source: Introducing Grok 4.5, SpaceXAI. Competitor scores are from published system cards or leaderboards, run with each provider's harness.

SWE-Bench Pro on the same post puts Grok 4.5 at 64.7%. Fable max leads at 80.4%. Opus 4.8 max is at 69.2%. Opus 4.7 max is at 64.3%, a hair behind Grok. GPT 5.5 xhigh is at 58.6%. That is a coding model in the cluster, at a much lower token price.

Token use is the other number they lead with. On SWE-Bench Pro, Grok 4.5 averages 15,954 output tokens per task, about 4.2× fewer than Opus 4.8 max at 67,020.

Composer stayed selectable. Grok 4.5 became the one the team reached for.

After the close, the family kept moving

Grok 4.6 shipped into the close. On CursorBench 3.2, scored with the Cursor agent harness, Grok 4.6 high is 69.9%, against 66.7% for Grok 4.5 high. DeepSWE v1.1 on that launch table is 65.9% versus 54% for 4.5 high. The Artificial Analysis Intelligence Index moves from 56 to 61, tied there with GPT-5.6 Sol Max. The price stays $2 / $6.

Grok 4.7, released September 21, is the model after the close: a larger base, a longer reinforcement-learning run weighted toward tasks that take many hours, served at the same price and speed as 4.6. CursorBench 4.0 is a harder, longer set, so its scores sit on a different scale from 3.2. On that new bench, Grok 4.7 xHigh lands at 46.3%, around $6 a task, ahead of Grok 4.6 High at 40.4% and GPT-5.6 Sol Max at 41.7%. Fable 5.1 Max is still ahead on the raw score, and it sits much further up the cost axis.

CursorBench 4.0 score versus average cost per task, from SpaceXAI's Grok 4.7 announcement

Source: Introducing Grok 4.7, SpaceXAI.

The comparison table is the within-family view. DeepSWE v1.1 moves from 65.2% on Grok 4.6 High to 71.0% on Grok 4.7 at high effort. Terminal-Bench 4.0 moves from 20.3% to 38.0%. The token rates stay put.

Grok 4.7 xHigh compared with Grok 4.6 High, GPT-5.6 Sol Max, and Fable 5.1 Max

Source: Introducing Grok 4.7, SpaceXAI. The asterisk on DeepSWE marks high effort.

A fuller split between "Cursor's loop got better" and "the weights got better" is in if it was only the harness. The short version: same-bench gaps inside the Grok family are model gaps.

What to change in review

A coding agent is trying to finish the task. A reviewer is trying to catch the change that should not merge. Long-horizon coding, tighter self-checks, and a flat $2 / $6 price are the traits that show up on a hard pull request.

In ThinkReview, Grok 4.5, Grok 4.6, and Grok 4.7 are selectable on their own, on GitHub, GitLab, Azure DevOps, and Bitbucket. The editor can stay Cursor. The review model is whichever one you pick in Model selection. OpenAI's November 12 wind-down of GPT inside Cursor is a separate contract. It leaves this path alone.

Use Grok 4.7 when you want the current coding default. Keep a second model on the diffs where you want a disagreement. The merge stays with the person who owns it.


Charts and scores reference Grok 4.5, Grok 4.6, and Grok 4.7 from SpaceXAI, plus Cursor's close-day post. Ready to run the current Grok on a real diff? Install ThinkReview or open Model selection.