If it was only the harness, Grok 4.5 would still be winning.

If it was only the harness, Grok 4.5 would still be winning.
Cursor's loop is good. It does not explain a model beating its own predecessor on a fixed bench.
When Grok started feeling sharper inside Cursor, two stories arrived together.
One: SpaceX bought the editor, pointed Colossus at the weights, and the model improved. The other: the weights were a sideshow. Cursor's harness — tools, repo context, terminal, and the room to iterate — is what makes a model look sharp.
Both describe a real afternoon in the product. They are different claims. A harness changes what a model is allowed to do. A training run changes what it can do when the harness is held still. The charts SpaceXAI posted are enough to separate them, if you refuse to mix benches.
The product history behind the jump is in Composer had the editor. Grok took the scores. This piece is only the measurement question.
Where the harness is the score
SpaceXAI says this out loud, which is useful.
CursorBench evaluates Grok with the Cursor agent harness on a fixed task set. The Grok 4.6 model card is explicit: 69.9% at high effort, and 70.8% at xhigh, is Grok inside that loop. Take the same weights into a thinner agent and the number should move. Anyone who says Grok "feels smarter in Cursor than in a raw completion" is often describing this, and they are right about the feeling.
Grok 4.5's own DeepSWE charts show the same trap with the labels still on. DeepSWE 1.0 was run by Artificial Analysis with each provider's harness. Grok 4.5 scores 62.0% there. DeepSWE 1.1, on the same announcement, is a mini-swe-agent run by Datacurve, and Grok 4.5 scores 53%. One model family, two loops, nine points. That gap is the harness.

Source: Introducing Grok 4.5, SpaceXAI. The footnote on the chart is the harness warning: each lab's own loop, via Artificial Analysis.
CursorBench 4.0 is a different test from CursorBench 3.2. SpaceXAI says 4.0 adds longer-horizon tasks, and that the scores are on a new scale. Putting 46.3% beside 69.9% and calling it a regression mixes two benchmarks. The harness changed jobs. The comparison broke.
Where the model is what moved
The comparison that holds is the one that keeps the loop fixed and changes the checkpoint.
On CursorBench 3.2, Cursor's harness, high effort: Grok 4.5 at 66.7%, Grok 4.6 at 69.9%. The model card's xhigh row for 4.6 is 70.8%, around 41,000 output tokens per task. The harness name did not change between those rows. The weights did.
DeepSWE v1.1 on the Grok 4.6 launch table: 54% for Grok 4.5 High, 65.9% for Grok 4.6 High. That is a different shape of result from a tooling tweak.
Then hold the later table still. The Grok 4.7 announcement puts 4.7 xHigh next to 4.6 High on one grid:
- CursorBench 4.0 — 40.4% to 46.3%
- DeepSWE v1.1 — 65.2% to 71.0% at high effort
- Terminal-Bench 4.0 — 20.3% to 38.0%
- Token price — still $2 in and $6 out

Source: Introducing Grok 4.7, SpaceXAI. Read 4.6 against 4.7 on this table. Read 4.5 against 4.6 on the 4.6 launch table. The DeepSWE figure for 4.6 is 65.9% there and 65.2% here; both are published, and they are different rows.
If the harness were the only upgrade since Grok 4.5, Grok 4.5 on that harness would still be the ceiling. Later Groks, scored the way SpaceXAI scores them against the previous Grok, are ahead of it.
The acquisition is what fed those runs. April was the Colossus partnership, after Cursor said Composer had improved every time the training run got bigger. June was the $60 billion agreement. August 14 was the close. Grok 4.7's launch post describes the part a harness cannot fake: a larger base, and a longer reinforcement-learning run on problems that take many hours. They also say 4.7 was trained to understand the Grok Bot harness. That is the honest middle. They fit the model to a loop, and they still publish same-loop deltas so the model can be seen moving on its own.
What transfers to a pull request
ThinkReview is a different harness from Cursor's agent. It reads the diff, pulls repository context, and writes findings a reviewer can act on. A CursorBench score does not copy over just because Grok feels strong in the editor.
What does transfer is the within-family lift. Better self-checks and longer tasks are exactly what a messy pull request asks for, including outside Cursor. Grok 4.7 is the checkpoint to select for that pass. To see the harness question on your own code, run the same diff on Grok 4.5 and Grok 4.7. Same review loop. Different weights. That is the comparison the model cards are describing.
Pick them in Model selection. The hosts are GitHub, GitLab, Azure DevOps, and Bitbucket. Treat the severity labels as input. Keep the merge with the person who owns it.
Benchmark details reference Grok 4.5, the Grok 4.6 model card, and Grok 4.7. Run both checkpoints on one diff with ThinkReview.