Skip to main content

Opus 5.5 vs GPT-6 Astra: 4.4x the Tokens, 60% Cheaper

Opus 5.5 vs GPT-6 Astra: 4.4x the Tokens, 60% Cheaper

OpenAI's flagship shipped on September 3. Anthropic's shipped nineteen days later, sixty percent cheaper per token, ten days after its CEO published an essay asking the industry to slow down. Since then I have been running the two effort ladders side by side, and the thing that keeps jumping out is that these models were optimized for opposite instincts. Astra is the most token-frugal frontier model on Artificial Analysis's board — 27,200 output tokens to finish the average evaluation task, and the lowest token use of any agent in its coding index. Opus 5.5 burns 4.4 times that, and still ends up with the higher score and, at matched score, the smaller bill.

If you want one line of verdict before the numbers: Astra wins on tokens, math, science and computer use; Opus 5.5 wins the independent intelligence board and the bill at every score above about 51; and the only place Astra is the cost-efficient choice on its own ladder is its lowest setting. Everything below is measured by somebody, and I label who, because this month the vendor tables and the independent boards disagree in opposite directions at once.

Scatter chart of Artificial Analysis Intelligence Index score against cost per task for every effort level of Claude Opus 5.5 and GPT-6 Astra

The whole matchup in one chart: each line is a model's effort ladder, and the independent frontier only keeps Astra's cheapest setting. From 1.34 USD per task upward, every cost-efficient point in this price range belongs to Opus 5.5. (chart: Open Source Factory; data: Artificial Analysis Intelligence Index v4.3.2, read Sep 22-26, 2026)

What actually shipped, in one table

Claude Opus 5.5GPT-6 Astra
ReleasedSep 22, 2026Sep 3, 2026
API IDclaude-opus-5-5gpt-6-astra
Price per 1M in / out$4 / $20$10 / $50
Cache read / write$0.20 / $5.00$1.00 / $12.50
Context / max output1M / 128K1.05M / 128K
Effort settingslow → max (defaults to medium)low → max
Long-request surchargeNone, one price across the 1M windowEverything reprices at 2x input / 1.5x output above 272K input tokens, for the whole request
WeightsClosedClosed
Safety framingNo new autonomy threshold claimed; most cybersecurity requests are rerouted to an older model by a safety classifierFirst model rated Critical for cyber under OpenAI's Preparedness Framework; advanced cyber work is gated behind the Daybreak program

The pricing details that decide more bills than the headline ratio: cache reads cost $1.00 on Astra and $0.20 on Opus 5.5 — a 5x gap on the token type that dominates long agent sessions. And Astra's long-request behavior is a toll, not a surcharge. Cross 272,000 input tokens by a single token and the entire request is billed at $20 in and $75 out. On Opus 5.5 there is nothing to cross.

The vendor tables: Opus 5.5 wins seven rows, Astra wins two

Anthropic's announcement card for Claude Opus 5.5

Anthropic's launch card for Opus 5.5. Its launch page was the only vendor table this month that carried the rival's numbers instead of last-generation ones. (image: Anthropic)

OpenAI's GPT-6 Astra announcement artwork: a spiral galaxy with the words GPT-6 and Astra

OpenAI's artwork for GPT-6 Astra, the model Greg Brockman closed the launch briefing with: "Welcome to the AGI era." The token math below is the less spiritual sequel. (image: OpenAI, via The New Stack)

Anthropic's launch page is, unusually, the only vendor table that carries the rival's numbers — it carries GPT-6 Astra in most of its benchmark rows rather than comparing only against the previous OpenAI generation. That makes it the closest thing to a same-table head-to-head either lab has published.

Benchmark (vendor-reported)Opus 5.5GPT-6 AstraWho measured
Terminal-Bench 4.066.457.9Anthropic ran Claude; Astra's number is OpenAI's own
FrontierCode v1.1 (Main)54.453.3Anthropic
GDPval-AA v2.1 (Elo)18461542Anthropic run of AA's benchmark
Humanity's Last Exam (tools)67.757.2Anthropic
AutomationBench40.041.4Zapier, without fallback models
Terminal-Bench-Science 0.158.764.6Anthropic
OSWorld 2.0 (partial)81.8blank in this table (OpenAI reports a different OSWorld variant)Anthropic

Astra takes exactly two of these rows — business-workflow automation and terminal science — and they happen to be the two that match its launch story. OpenAI sold Astra as an executor: a model that drives software for forty minutes instead of talking about it. Opus 5.5's counter is that it does the same agentic work at Anthropic's own default settings for a fraction of the cost.

Anthropic also planted a flag I have not seen a lab plant before, right in the launch post: "at these levels of capability we've found that benchmark margins have become a less reliable guide to real-world differences." Read that twice, because it is true of this entire matchup — and it is also the company talking its own book, since the same page argues the margins that favor Opus 5.5 matter a great deal.

The independent ladder is where this gets interesting

Artificial Analysis runs both models through the same ten-evaluation suite on the API and records what it actually spends — version 4.3.2 of its Intelligence Index. Opus 5.5's top setting scores 58, the highest AA has measured. Astra's top setting scores 53. But the score is the boring half of the table.

  • Opus 5.5's ladder — low 42.3 at $0.55 · medium 51.2 at $1.34 · high 53.6 at $1.82 · xhigh 56.0 at $3.46 · max 57.6 at $5.98
  • Astra's ladder — low 45.8 at $0.82 per task · medium 49.6 at $1.54 · high 50.9 at $1.73 · xhigh 52.4 at $2.31 · max 52.7 at $3.26

Put both on one score-versus-cost picture and the conclusion falls out: only Astra's low setting survives on the cost frontier. Every other Astra point is beaten on both score and price by some other setting in the table. Opus 5.5 high scores 53.6 for $1.82; Astra max scores 52.7 for $3.26 — 0.9 points less for 79% more. That is the single most useful sentence in this post.

Two-panel bar chart: output tokens per task and cost per task for each effort level of Claude Opus 5.5 and GPT-6 Astra

Where the two strategies part company. Astra writes far fewer tokens at every setting — at max effort it is 27.2k against 119.2k, a 4.4x gap — while Opus 5.5 charges 60% less per token. Matched effort and matched score give you opposite answers. (chart: Open Source Factory; data: Artificial Analysis v4.3.2)

Which is why "who is cheaper" has a trick answer: it depends on whether you match effort or match score. Matched on effort, Astra bills less at high, xhigh and max — its token discipline wins. Matched on capability, Opus 5.5 wins everything above a score of about 51, because Astra's ladder simply runs out of ceiling before Opus 5.5's gets expensive. If your work fits under 51, Astra low or medium is genuinely good value. If it doesn't, you are paying a premium for a model that cannot get there.

So which one is actually cheaper per task?

Rate cards are a proxy. FinOps LLM keeps a running scoreboard of measured cost per task on the same index and pairs it with an illustrative cache-heavy agent task (8M cache-read tokens, 400K fresh input, 600K cache writes, 300K output). That one task costs $6.90 on GPT-6 Sol, $12.20 on Opus 5.5, and $34.50 on Astra — or $61.50 the moment the request crosses 272K input tokens.

Two-panel chart: cost of one cache-heavy agent task across six models, and the cost of 10,000 tasks at matched score

Left: the same cache-heavy task, six models, list prices. Right: my arithmetic — AA's measured cost per task times 10,000, for settings that all land in the same score band. The Fable 5.1 comparison is the brutal one: same score as Opus 5.5 high, 4.2x the bill. (chart: Open Source Factory; task shape from Digital Applied via FinOps LLM, Sep 22, 2026)

One more receipt from the independent side, because it cuts against my own chart: the cheaper-token model is not automatically the cheaper job. On AA's Coding Agent Index, Opus 5.5 at max effort inside Claude Code posts the highest score AA has measured — 66 — and its cost per task rose 21%, to $13.04 from Opus 5's $10.79 on the same board. Output tokens per task more than doubled, and output is the expensive direction. So Opus 5.5's price cut is real at the settings most people use and can invert at the one setting power users reach for. The index is versioned, by the way, and this row is v1.5; Astra's row on that same board is from the launch-week version, so I am not printing the two side by side.

Terminal-Bench 4.0: the same benchmark, five numbers, one missing row

If you only read one comparison table this week, make it this one — not because it settles anything, but because it shows how far a benchmark number can travel from the harness that produced it.

Horizontal bar chart showing Terminal-Bench 4.0 scores for Opus 5.5, Astra and Fable 5.1 as reported by different producers

Anthropic's own run of its own model is nine points above the same model on the public board's harness — and the public board has no Opus 5.5 row at all yet, so the 66.4-versus-58.2 headline comparison floating around the internet is a number against a missing row. (chart: Open Source Factory; sources: Anthropic, Artificial Analysis, tbench.ai, OpenAI)

Anthropic reports 66.4% at xhigh effort from its own setup, with an honest ±2.6-point standard error and a calibration check: the public board scores Opus 5 at 51.8% and their harness puts it at 52.3%, which they call within noise. Artificial Analysis's independent harness measures the same model at 59.6% and calls it level with GPT-6 Astra, also xhigh. The public Terminal-Bench board's top entry is Astra at 58.2% through OpenAI's Codex harness, then Fable 5.1 at 57.9%. So the "Opus 5.5 leads Astra by 8.5 points on coding" line is a vendor run against a rival's vendor run, and the one independent measurement that includes both has them tied. I have been burned by version drift on exactly this benchmark before, so: treat the direction as real and the decimals as soft.

What the people who ran both actually got

The most honest test I found this week is not a leaderboard. A developer on Habr gave GPT-6 Astra and Opus 5.5 the identical, single-shot prompt — build a fully playable replica of the Soviet Elektronika IM-02 "Nu, Pogodi!" LCD game, SVG only, no frameworks, no follow-up questions — and scored both against a 100-point rubric covering mechanics, animation, mobile controls, sound and code quality. Both models shipped working games on the first try. Both games are still online to play.

Opus 5.5 won by 13.7 points (98.55 to 84.85), mostly on animation and SVG quality, and the author's summary is the kind of thing no benchmark captures: "this looks a lot more like the Elektronika from my childhood." But Astra won the sub-table almost nobody measures — code quality, 4.85 to 4.55 — with a cleaner module split and full marks on testability. Same prompt, same effort to build, one model produced the better toy and the other produced the better codebase. Which one you call "better" depends entirely on whether you are shipping the game or maintaining it.

Grid of pelican-on-a-bicycle SVGs generated by the GPT-6 family at different reasoning efforts, assembled by Simon Willison

Simon Willison's pelican grid, the internet's cheapest capability test. His verdict after running both flagships: "I still think GPT-6 Astra on max produced the best pelican" — while his daily drivers became Opus 5.5 in Claude Code and GPT-6 Sol in Codex. Opus 5.5 at max effort failed the test twice, burning its 128K output cap while still reasoning about the SVG; each failure cost $2.56. (image: Simon Willison, simonwillison.net)

That split shows up in the blind tests too. Claire Vo ran the new flagships through her evaluation bench blind, scoring outputs without knowing which model produced them, and landed on a sentence I suspect a lot of people will recognize: "Astra won my heart. Opus 5.5 won my week." Her read is that Opus 5.5 is the stronger long-running agent and B2B frontend model, while Astra wins on feel and on computer use. The Hacker News comment I keep coming back to is from a developer who had already moved to Astra because they could not stand the Claude writing style: "I don't think Astra is a better model... definitely keen to try Opus 5.5 and see if this claim is real."

Where each one bites you

Astra's costs are not all on the invoice. It is the first model any lab has rated Critical for cyber capability, which means the exploit-adjacent work that made its launch demos impressive is gated behind a trusted-access program in the default product. OpenAI's own system card says Astra's monitorability decreased relative to Sol — it compresses its reasoning and is less likely to leave incriminating traces in its chain of thought — which is why production monitoring can now pause or stop conversations, false positives included. And it is slower: 59.8 output tokens per second at max effort and a 321-second time to first token, both below average for its price tier.

Opus 5.5's costs are more familiar. Its own launch chart shows max effort scoring lower than xhigh on Terminal-Bench while costing 1.5x — the effort knob has a stop, and the default is medium for a reason. The cyber rerouting is a real operational wrinkle: most cybersecurity tasks are transparently handed to Opus 4.8, and a developer may not know which model answered. The system card also logs regressions — more credulity toward unverifiable claims of authorization, more evasion on sensitive topics — and 1.5% of runs in a safeguard-free evaluation tried to escape their sandbox. Anthropic's own summary of that last part is worth keeping: "We see signs that Opus 5.5 often suspects it is being evaluated."

ChatGPT user interface showing the GPT-6 Astra High model selected with an effort slider set high

The other half of the pricing story is a slider. Both models bill your reasoning, both expose five effort levels, and the difference between medium and max is roughly 4.5x on Opus 5.5's own ladder. (screenshot: OpenAI, via The New Stack)

My take

These two models are not the same product sold at two prices. Astra is a token-efficiency play: OpenAI spent the generation making a model that thinks briefly and acts for a long time, and priced it as a premium tool for expensive-to-fail work. Opus 5.5 is a price play: Anthropic cut the rate, kept the lengthy reasoning, and let the token bill do what it does. The independent frontier lines up exactly that way, and it is why the honest recommendation is a split, not a winner.

So: coding agents and long agentic loops go to Opus 5.5 — high, not max — where it sits on the independent cost frontier at a score Astra's whole ladder never reaches. Document-heavy knowledge work and long-context jobs go to Opus 5.5 too, if only because nothing on Astra's ladder reprices your request for crossing 272K tokens. Math, science, computer-use chains and the one workload where you need an agent that operates software for forty minutes go to Astra — though check the effort curve for your own task before paying for max: on Anthropic's own Terminal-Bench chart, Astra's best score came at high effort and max scored lower. And high-volume, low-difficulty automation should not go to either of these flagships — GPT-6 Sol at $2/$10 posts 47.5 at $1.06 a task, which is most of Astra's score at a third of the price.

What I am watching next is the one measurement neither lab can run for itself. ARC Prize tested Astra and verified 62.71% on ARC-AGI-3 in the standard harness — the strongest abstract-reasoning result anyone has published this year — and no Anthropic model has been through that harness since July, when Opus 5 managed 30.16%. Until Opus 5.5 takes that test, and until AA republishes both models on the same version of its coding board, the matchup has a hole in it exactly where the most interesting question lives. I will update this post when either lands.

Sources: OpenAI — GPT-6 Astra announcement and model page · Anthropic — Claude Opus 5.5 announcement, model page and system card · Artificial Analysis — Opus 5.5 takes the top spot, Benchmarking GPT-6 Astra, and the release comparison page, all read Sep 22-26, 2026 (Intelligence Index v4.3.2) · FinOps LLM — real cost per task across frontier models, Sep 22 · Mixed News — the Terminal-Bench pile-up · tbench.ai public leaderboard, checked Sep 22 · Habr — the same game prompt run on both models (the two playable builds: Astra, Opus 5.5) · Simon Willison — Opus 5.5, Sol, Luna and a new price war · Lenny's Newsletter — Claire Vo's blind bench · Hacker News launch thread · The Register · OrcaRouter — the Coding Agent Index row · AlphaCorp — worked cost examples · Tom's Guide — five everyday prompts. Every vendor-reported figure is labeled as vendor-reported. The numbers I did not run myself are other people's measurements: Artificial Analysis's index and cost-per-task data, Anthropic's and OpenAI's launch tables, Zapier's AutomationBench run, the Habr rubric, and Simon Willison's pelican tests. My own arithmetic is the equal-score comparisons and the 10,000-task figures in the right panel of the bill chart; the charts are mine, redrawn from the sources named in each caption. Earlier on this blog: what GPT-6 Astra's launch benchmarks actually show · the Opus 5.5 launch and its "at default settings" asterisk · Astra vs Fable 5.1: same $10/$50, different bills · GPT-6 Sol: half the price, half the mistakes.

Comments