Skip to main content

Grok 4.7: Terminal-Bench 20% → 38%, Same $2/$6 Price

Grok 4.7: Terminal-Bench 20% → 38%, Same $2/$6 Price

SpaceXAI is the AI division of SpaceX — logo: SpaceXAI, public domain via Wikimedia Commons

Musk said ten days. It took nineteen.

On September 2 he wrote that Grok 4.7 would land in ten days, which pointed at September 12. It shipped on September 21 — forty days after Grok 4.6, nine days after his own target, and a week after he'd already started talking about Grok 4.8 instead. So the interesting question isn't whether the wait was worth it. It's what the delay did to the numbers.

Here's the short answer, and I'll show the receipts below: Grok 4.7 is a same-price upgrade that nearly doubles xAI's long-horizon coding score. It is not the best model in the world, and xAI's own launch card quietly shows you why.

What actually shipped

Grok 4.7 is a new, larger base model than Grok 4.6, trained with a longer reinforcement-learning run on a harder task mix — weighted, in xAI's words, "toward problems that take many hours to complete." The launch post also says the model is better at verifying its own work and at managing long context, and that it was trained to natively understand the Grok Bot harness, which is the piece that matters if you're one of the people running always-on Grok Bots rather than chatting.

Pricing is the part that will get quoted: $2 per million input tokens and $6 per million output tokens, unchanged from Grok 4.6, with a fast variant at twice the output speed for twice the price. Context stays at 500K. Weights are still proprietary, and xAI still hasn't published a parameter count — the widely repeated "2.1 trillion" number comes from Musk's September 2 posts on X, not from the launch page.

Grok 4.7 at a glanceDetail
Base modelNew and larger than Grok 4.6. No official parameter count (the 2.1T figure is Musk's)
Context window500K tokens, unchanged
List price$2 per 1M input tokens / $6 per 1M output tokens — same as Grok 4.6
Fast variantTwice the output speed at twice the price
Where to run itCursor, Grok Build, xAI API, third-party harnesses, routers and cloud platforms
Trained forLong RL run weighted toward multi-hour tasks, self-verification, and the Grok Bot harness
Safety claimsNew safeguard stack; 62.4% on LatchBio's biosafety benchmark; 3.3% risky-prompt pass-through on HackerBench v0.3 — all vendor-reported
Grok 4.7 benchmark comparison card from xAI

xAI's official launch card compares Grok 4.7 xHigh against Grok 4.6 High, GPT-5.6 Sol Max and Fable 5.1 Max. Every figure on it is vendor-reported — image: SpaceXAI

Read that card the way a competitor would. Grok 4.7 wins five of the seven benchmark rows against GPT-5.6 Sol and loses two (DeepSWE v1.1 71.0% against 72.7%, HealthBench Professional 56.7% against 60.5%). Against Fable 5.1 it takes three rows and drops four. It beats both of them badly on price, though: Sol charges $4 in and $20 out, Fable 5.1 charges $10 and $50.

Percentage-point gains of Grok 4.7 over Grok 4.6 across six benchmarks

Every row on the card moved up from Grok 4.6 at the same $2/$6 price. Terminal-Bench 4.0 gained 17.7 points — my chart, from xAI's vendor-reported numbers

Terminal-Bench 4.0 going from 20.3% to 38.0% is the number to remember. That's the benchmark where Grok 4.6 got hammered in August — 26% on the 3.0 version of the suite, against GPT-5.6 Sol's 34.6% — and it's the one closest to what agentic coding actually feels like: a model left alone in a terminal, tasked with finishing something multi-step without a human nudging it. Nearly doubling it in forty days is the single most interesting thing on the page.

The delay was the story

Most release posts skip the awkward part, so let me not. On September 11, a day before the promised date, Musk explained the holdup in a reply on X, and the reason was more specific than "polishing":

"Grok 4.7 needs a few more days to cook. We might have penalized response length too much (or something) in RL, as it still gives up on hard tasks (that it can do!) too early and isn't yet sufficiently rigorous in checking its work."Elon Musk, on X, September 11, 2026 — 2.1M views

Now put that next to the launch post, which says Grok 4.7 "works longer on difficult tasks" and "checks its own work more carefully." The two sentences are the same sentence — one is the bug report, the other is the changelog. And that's the cleanest evidence the extra nine days went somewhere real, rather than into a marketing calendar.

Timeline from Grok 4.6 in August to Grok 4.7 on September 21

Forty days from 4.6 to 4.7 — with Musk naming Grok 4.8 before 4.7 had shipped. My chart, dates from xAI's news posts and Musk's X replies

Meanwhile the rest of the industry wasn't waiting. The first three days of September alone produced Claude Fable 5.1 and Mythos 5.1, Meta's Muse Spark 1.3, Gemini 3.8 Flash, and GPT-6 Astra. Grok 4.7 arrived three weeks into that wave, and by then Musk had already recalibrated expectations in public — "roughly on par with Opus 5.0, not 5.1," he said on September 14, "better in some ways, worse in others," pointing at multimodal performance as the gap that still needed work.

Elon Musk at the U.S. Air Force Academy

Musk spent the delay posting the roadmap instead of the model: 4.8 finishing pretraining, 4.9 "probably Astra/Fable class," Grok 5 as the AGI bet. Photo: U.S. Air Force / Trevor Cokley — public domain, via Wikimedia Commons

The graph everyone noticed

The launch card's four columns are Grok 4.7, Grok 4.6, GPT-5.6 Sol and Fable 5.1. Eighteen days earlier OpenAI shipped GPT-6 Astra, and it appears nowhere on the card. The Hacker News thread on the announcement zeroed in on that fast — one commenter asked whether leaving Astra out "can't have been an oversight," and a reply gave the answer that holds up:

"You can't use Astra in Cursor, and cursorbench uses cursor as the harness. They can't actually benchmark it using their harness hence why its not included."A reply in the Hacker News thread on the Grok 4.7 announcement

That's not a dodge, it's a consequence of the corporate situation. On August 28 OpenAI said it would wind down the contract that supplies its models to Cursor — shutoff proposed for November 12 — and that it would stop providing future models to Cursor, naming Astra as the model it wants held to its terms. Astra launched six days later. A CursorBench 4.0 table physically cannot contain it.

You can still read the omission as a choice, because the card's ordering is xAI's to make, and the rival they did put in the second OpenAI column is the one Grok 4.7 beats. That's normal launch-page behavior, and it's why I don't treat vendor cards as scoreboards — I treat them as a specification of what the vendor thinks it can win.

Third parties close the gap. Cursor's own public CursorBench leaderboard is still on version 3.2 and doesn't list Grok 4.7 at all yet, so the 46.3% on the card is currently uncross-checkable. Benchmark aggregator BenchLM marks Grok 4.7 "not publicly ranked," with the best verified rows belonging to Fable 5.1 on CursorBench (51.8%), Muse Spark 1.3 on DeepSWE (75.4%), Mythos 5.1 on Terminal-Bench 4.0 (60.9%) and GPT-6 Astra on HealthBench Professional (63.4%).

So what does the independent read say?

Artificial Analysis, which measures models itself rather than reprinting vendor numbers, has Grok 4.7 at 46 on its Intelligence Index. That is two points ahead of Grok 4.6 (44) and one point behind GPT-5.6 Sol (47) — a very small step, honestly, for a model that's larger than its predecessor and a full month newer.

Artificial Analysis Intelligence Index scores for Grok 4.7 and rival models

Where Grok 4.7 lands on Artificial Analysis' index, with each model's list output price. Grok 4.7 is the cheapest model on this list except Gemini 3.8 Flash and Muse Spark 1.3 — my chart, data from artificialanalysis.ai, retrieved September 22

Look at the orange bar and then look at the row under it. Muse Spark 1.3 scores 48 at $1.25 in and $4.25 out. Gemini 3.8 Flash scores 41 at $0.75 and $3.75. Grok 4.7's real claim isn't "frontier for cheap" — it's "frontier-adjacent at a price the frontier stopped charging", and it's sharing that shelf with a couple of cheaper models. The top of the list is still Astra and Fable 5.1 at 53, and both of them cost $10 in and $50 out to get there.

There's a second, less flattering number hiding in Artificial Analysis' data. To complete their Intelligence Index, Grok 4.7 generated 240M output tokens. Grok 4.6 used 94M for its run. AA labels that "very verbose" — 4 out of 4 on their verbosity scale. Caveat, and it's a real one: 4.7 was run at its xhigh reasoning effort and 4.6 at high, so some of that gap is a setting, not a personality trait. But a commenter in the launch thread framed the underlying issue better than any chart does — "Token price doesn't tell you much without knowing token efficiency." If your agent burns 2.5x the output tokens per task, "same price" doesn't mean "same bill."

Two data points AA hasn't published yet: 4.7's speed and its cost per Intelligence Index task. Both were still blank when I pulled the page. I'll update this post when they land.

What it costs, and where you can run it

List price is the same as Grok 4.6: $2 in, $6 out, with a fast variant at $4/$12. The interesting wrinkle is what the routers show. On OpenRouter, xAI's own endpoint lists Grok 4.7 at $1.60 in and $4.80 out with cache reads at $0.40 — 20% under the list price, and under Grok 4.6's $2/$6 on the same page. xAI's blog post doesn't mention a discount, so I'm not going to invent a reason for it. Either it's a launch promotion or a quiet cut that hasn't reached the announcement; both are worth knowing before you commit a codebase to it.

Availability on day one: Cursor, Grok Build, and the xAI API — plus third-party coding harnesses, model routers and cloud platforms. There's no mention of grok.com or the X apps in the announcement, which is consistent with Grok 4.6's pattern: the coding surfaces get the model first, the chat apps catch up later. If you want to poke at it without paying, xAI is pointing people at Grok Build (x.ai/build).

Grok Build agent editing a TypeScript file in the terminal UI

Grok Build is where xAI wants you to try 4.7 first — the agent harness the model was trained to understand natively. Image: SpaceXAI

Is Grok 4.7 better than Fable 5.1?

No — not on the benchmarks, and xAI's own card is the evidence. Fable 5.1 wins CursorBench 4.0 (51.8% vs 46.3%), Terminal-Bench 4.0 (57.9% vs 38.0%), HealthBench Professional (62.1% vs 56.7%) and AA Briefcase v1.1 (1,678 vs 1,657). Grok 4.7 wins EEBench (64.0% vs 56.4%) and the legal-agent benchmark by an absurd margin (19.6% vs 6.7%). Fable 5.1 also costs five times more on input and more than eight times more on output.

So the honest framing is a question of workload, not of ranking. If you're paying per token for long agentic runs, Grok 4.7's price-performance is genuinely hard to argue with — that's what "frontier in price-performance" is supposed to mean. If you need the top score and someone else is paying, it isn't the model you buy.

The community split tracks that. In the same Hacker News thread, one commenter who had been running Grok in Cursor for months reported the results were "highly superior" to another agent, and another called Grok Build "a no nonsense model and stays on its course." A third described the opposite experience — Grok ending tasks almost immediately and claiming "Done!", calling it "the laziest and most 'dishonest' of all the models... I can't use Grok for any serious coding task." Neither comment is about 4.7 itself — nobody in that thread had run it yet — and the second describes the exact failure mode Musk said the RL run had caused. That's the thing to test in your own repo this week, not in anyone's benchmark table.

My take

This is my third Grok write-up in six weeks, and the first where a delay did something for me. Grok 4.6 was a same-size refresh with better training — impressive, but it didn't change what the model was for. Grok 4.7 looks like it was pointed at the specific thing that made Grok hard to trust on long tasks: quitting early and not double-checking. The Terminal-Bench jump and the self-verification claims are the same fix, described from two directions.

What I'm not going to pretend: one vendor card, one independent index that shows a two-point gain, and a public safety claim ("strongest model we've tested on refusals and jailbreak resistance", 62.4% on LatchBio's biosafety benchmark) that nobody outside xAI has verified yet. The card's omission of Astra is defensible for mechanical reasons and still worth remembering when you read the headline number.

Next on the calendar: Musk said on September 14 that Grok 4.8 — a 2.5T model on xAI's new C++ training stack — would finish pretraining that week and start RL, and that Grok 4.9 is "probably Astra/Fable class." If that pacing holds, this post has a short shelf life and I'll be writing the sequel soon. When Artificial Analysis publishes 4.7's speed and cost-per-task, I'll come back and update the numbers in this one.

Comments