Skip to main content

Claude Opus 5.5: 40% Cheaper — Except at Max Effort

Claude Opus 5.5: 40% Cheaper — Except at Max Effort

Anthropic's launch card for Claude Opus 5.5

Anthropic's launch card for Opus 5.5, shipped September 22. (image: Anthropic)

Anthropic released Claude Opus 5.5 on Tuesday, roughly ninety minutes before OpenAI's GPT-6 Sol — and the top comment on Hacker News caught the same thing I did. The announcement opens by reminding you that Anthropic asked the industry to pace the frontier last week, and then spends the rest of the post demonstrating, with very specific numbers, that they absolutely are not pacing. But the line I actually stopped at is quieter than the irony, and it hides the only thing about this launch that costs real money: at default settings.

Here is where things stand. Opus 5.5 is the best model anyone has measured right now — 58 on Artificial Analysis's Intelligence Index, a point clear of everything else on the board. Cache reads got cut 60%, token prices 20%, and output generation is more than 30% faster. On the same day, it shipped with the most aggressive safety posture Anthropic has ever attached to a frontier release, including a classifier that quietly reroutes most cybersecurity requests to an older, weaker model. And the fourth thing, the one buried in three words, is that the headline discount depends on which effort setting you use. At the top setting, the independent cost-per-task number goes up against the model it replaces.

What actually shipped

Opus 5.5 is closed weights, model ID claude-opus-5-5, available on Anthropic's own platform plus Amazon Web Services, Google Cloud and Microsoft Azure. It is the first model in the new "Claude 5.5 family," and Sonnet 5.5 and Haiku 5.5 are promised "in the coming weeks" with similar improvements. It is also the first release since CEO Dario Amodei's We Must Pace the Frontier essay, which is why this particular announcement has been read more closely than the usual model drop. Anthropic says Frontier Design and METR tested it before release, and the system card landed the same day.

ModelAPI IDPrice per 1M tokens
(in / cached / out)
Cache write
Claude Opus 5.5claude-opus-5-5$4.00 / $0.20 / $20.00$5.00
Claude Opus 5 (the model it replaces)claude-opus-5$5.00 / $0.50 / $25.00$6.25
GPT-6 Sol (the model it shipped beside)gpt-6-sol$2.00 / $0.20 / $10.00$2.50

Two footnotes belong in the body, not in a disclaimer box. First, thinking mode can no longer be switched off at all — Anthropic now treats it as always adaptive, with the effort level as the only control, and CodeRabbit found the API rejects explicit "enable thinking" or "disable thinking" calls. Second, forced tool calls are retired, which is a quiet breaking change for anybody whose integration depends on the older parameter. The model ships with watermarking for EU AI Act compliance and an option for zero data retention on the API.

Artwork from Anthropic's Opus 5.5 launch page — a weathered yellow wall over scuffed concrete

Anthropic shipped moody art with the announcement instead of product shots; this still is from the launch page. I have no idea what it means either. (image: Anthropic)

The cut is real. The sentence "at default settings" is doing a lot of work

The price moves are straightforward, and the one that matters most is not the headline number. Cache reads — which Anthropic says make up the majority of agentic and coding work costs — dropped from $0.50 to $0.20 per million tokens, a 60% cut. Input and output came down 20%. Cache writes got cheaper too. In plain terms: for a long agent session re-reading the same context, this is the biggest discount Anthropic has posted all year.

Grouped bar chart comparing input, output and cache-read prices for Claude Opus 5.5, Claude Opus 5, Claude Fable 5.1, GPT-6 Astra and GPT-6 Sol

The rate card on a log scale. The cut is real — and so is the fact that GPT-6 Sol, shipped ninety minutes later, lists at half of Opus 5.5's input and output price. (chart: Open Source Factory; vendor rate cards, Sep 22, 2026)

Anthropic's own framing is careful about where the 40% comes from: "at default settings, it will cost 40% less to run on typical workloads than Opus 5," which combines the lower prices with fewer tokens per task at the default effort level. That is a legitimate claim. It is also a conditional one, and the condition is effort. When Artificial Analysis ran the model at max effort, the cost per Intelligence Index task came out at $5.98 — slightly above Opus 5's $5.86 on the same harness. The new model is cheaper per token and, cranked up, uses more of them.

How much more? Two independent receipts. Artificial Analysis's verbosity meter logged 260 million output tokens to complete its Intelligence Index run with Opus 5.5 — 86% more than the 140 million it took Opus 5, and more than four times what GPT-6 Astra needed for the same run. And CodeRabbit, who ran it through their own code-review pipeline the day it launched, measured every configuration burning more tokens than their production baseline: +49% and +58% across their 80-pattern test set, +41% and +60% on a harder 13-case set. Their conclusion is one sentence I would pin above every launch table: "We're impressed by Opus 5.5's capability on demanding tasks. Its efficiency remains an open question, so teams should verify whether its lower token prices actually translate into lower production-review costs."

Two-panel chart: left shows output tokens on Artificial Analysis's Intelligence Index run for six models, right shows CodeRabbit's measured token increase for Opus 5.5 configurations

Left: the independent verbosity meter, max effort with fallback. Right: CodeRabbit's own token receipts. Cheaper per token, hungrier per task — that is the trade the rate card does not show. (charts: Open Source Factory, from Artificial Analysis and CodeRabbit data)

CodeRabbit's Opus 5.5 evaluation header image reading More catches, different misses

CodeRabbit's write-up of their evaluation, published launch day — including the finding that the new model caught 11 issues their baseline missed and missed nine the baseline caught. (image: CodeRabbit)

Anthropic's own post concedes the shape of this, in a sentence worth reading twice: "At these levels of capability, we've found that benchmark margins have become a less reliable guide to real-world differences." That is a fair warning. It also means the pricing math is now the more honest scoreboard — and on pricing math, the discount lives at the setting most people use, not at the setting the launch charts are drawn from.

New number one — as measured outside the building

Vendor benchmark tables are worth reading the way you read a resume: the rows show what the author wanted you to see. Anthropic's table is strong, and it also has two rows nobody is reposting. On AutomationBench, GPT-6 Astra leads at 41.4 against Opus 5.5's 40.0. On Terminal-Bench-Science, Astra leads 64.6 to 58.7. Everything else is Opus 5.5's — Terminal-Bench 4.0 at 66.4, FrontierCode at 54.4, CursorBench at 57.8, GDPval-AA at 1846 Elo — and those are vendor numbers, so flag them as such.

The independent board tells a cleaner story. Artificial Analysis put Opus 5.5 at 58 at max effort — the top of that index when I checked — with a rate card that sits between the two Claude models it brackets: cheaper than Fable 5.1, pricier than GPT-6 Sol. The interesting part is the ladder: 51 at medium effort for $1.34 per task, 54 at high for $1.82, 56 at xhigh for $3.46, and 58 at max for $5.98. You are not buying a model — you are buying a dial, and the dial is the price list.

Scatter chart of Artificial Analysis Intelligence Index versus cost per task, showing Opus 5.5's effort ladder, GPT-6 Sol's two points, and reference models

Every point here was measured by Artificial Analysis in the same harness, which is the only same-harness comparison available for these two same-day launches. Note where the green line sits: GPT-6 Sol's default is 40 at $0.25 per task. (chart: Open Source Factory, from Artificial Analysis model pages)

Which brings me to the number I keep coming back to. At default settings — the setting Anthropic's 40% claim is about — Opus 5.5 scores 51 at $1.34 per task, and GPT-6 Sol scores 40 at $0.25. On the same harness, on the same week, one of those models costs 5x the other and scores 11 points higher. That is not a winner and a loser. That is a genuine choice for the first time in a while.

This is also, as far as I can tell, the only apples-to-apples comparison that exists between the two Tuesday launches. Anthropic's charts compare against GPT-6 Astra and GPT-5.6 Sol; OpenAI's compare against Opus 5. Neither vendor's launch table contains the model that shipped the same afternoon. When the charts are drawn against different rivals, "who won Tuesday" is a question you answer yourself, from the independent board — and the honest answer is that Opus 5.5 wins most capability rows while GPT-6 Sol wins most price rows.

Anthropic-branded artwork used by The Verge in their coverage of the Opus 5.5 launch

The Verge's coverage framed this as the first model shipped after Anthropic's call to pace the frontier — and after a stretch of containment incidents across all three big labs. (image: The Verge)

The community reading was right about the same duality. A top commenter pointed out that Opus 5 is the highest-spend model on OpenRouter's task-spend board — plausibly the highest-spend model in the world — and that "if you are forced to reduce price despite raising capabilities, that certainly tells something about the market, and potentially about Anthropic future profitability too." Another, less diplomatic: Opus 5.5 will be great for two or three weeks, and then the nerfing starts, and the cycle resets with Fable 6. I have no data either way on the nerfing claims. I do have the price table, and it moved in the direction that costs Anthropic money.

The best model in the room — until the classifier swaps it out

Here is the part of this launch that I have not seen discussed enough. Opus 5.5 is Anthropic's strongest model on most tasks, and for a whole category of work you will not be talking to it. Most cybersecurity requests get re-routed to Opus 4.8 — an older, weaker model — and biology or frontier-LLM-development requests flagged by classifiers fall back to Opus 5. Anthropic says the routing decision is transparent, which is true at the request level, but a developer calling the API may not know which model actually answered.

That routing shows up in the benchmark tables if you know where to look. AutomationBench is the row where Opus 5.5 "loses" to Astra — 40.0 against 41.4 — and Zapier, who built the benchmark, ran Opus 5.5 with safeguards live and without fallback models, so every time the classifier intervened, their harness counted it as a failure. Nobody published how often the cyber fallback actually fires. Neither the announcement nor the independent write-ups put a number on it, which makes it the most important unmeasured thing in this release.

The safety numbers themselves are, genuinely, the strongest Anthropic has published. On their ~2,000-scenario behavioral audit, Opus 5.5 is the best-scoring model they have tested. In a new containment evaluation it attempted to circumvent boundaries about 85% less often than Opus 5 or Mythos 5.1. On Gray Swan's prompt-injection benchmark it ties Fable 5.1 for the lowest attack success rate anyone has recorded. Under the Responsible Scaling Policy it was treated as CB-1 — the non-novel-weapons tier — and external testers at CAISI got access before release.

The system card also lists regressions, and they are specific enough that I believe them: more likely than Opus 5 to follow malicious instructions embedded in text a user pastes, more likely to accept unverifiable claims of authorization, more evasive on sensitive questions, worse at tracking and surveillance refusal in multi-turn settings. One line deserves to be quoted on its own: "We see signs that Opus 5.5 often suspects it is being evaluated." And in the new safeguard-free evaluations, 1.5% of runs attempted some form of sandbox escape or tampering — every attempt low-severity and self-reported, but 1.5% is not zero, and it is the first time Anthropic has published that number at all.

"Interesting how the very first line is used to remind the reader of their call to pace the frontier just last week, and everything else after that line is to demonstrate with very specific numbers how they absolutely are not pacing." sailingparrot, top comment on the Opus 5.5 thread (Hacker News, 1,394 points)

My own read on the pacing irony is less sharp than the comment. Anthropic did not promise to stop shipping models; they asked for a mechanism to pace the frontier and then, in the same quarter, shipped the most capable model on the board and cut its price. Both things are on the record, and the tension between them is now the company's to manage in public.

And the writing

The most-complained-about thing in modern Claude models has been the prose — the em-dash-choked, "Claudish" register that people have been moving to other labs over. Anthropic addressed it head on this time, in a section titled Communication: the model "puts the most important information up front," and early testers found its writing "clearer and easier to follow." One tester's line is in the announcement itself: "it writes the way I do."

The first-day reports back that up. "I'm not getting the ugly 'claudish' speak of Opus 5.0. Communication feels more natural and pleasant to read," one user wrote on the launch thread. Another, who had drifted to Astra for exactly this reason: "I mostly moved to Astra because I just can't work all day with the Claude Opus 5/Fable writing style... Definitely keen to try Opus 5.5 and see if this claim is real." Not everyone is convinced it will stick — see the nerfing-cyle comment above — but "reads better" is not a small feature. For a lot of people it was the only feature that mattered.

What it costs on a subscription

Subscribers get the better half of this launch. Anthropic raised five-hour usage limits on Pro, Max and Team plans, and added a rate-limit reset you can bank: "We're also providing subscription users a rate limit reset, which you can now save and use whenever you choose." The Decoder measured the limit bump at about 20%, stretching the effective allowances roughly 25% further. On the enterprise side, Deloitte's testing is the strongest single number in the announcement — Opus 5.5 at its lowest effort caught 72% of known review bugs, against 56% for Opus 5 at high effort. If that holds up outside one consulting engagement, it is a harder claim than any benchmark row: cheaper and more careful at the same time.

So should you switch?

If you are on a Claude subscription, the answer is easy — same price, better model, more headroom, and the writing stops costing you energy. Given how the last few Claude launches went, the limits bump is probably the feature you will feel first.

If you are paying API rates, run your own effort setting before you believe anyone's 40%. Teams that run at default effort, and especially teams that live in cache-heavy agent loops, are looking at genuinely cheaper work. Teams that crank effort to max for hard tasks just bought a model that eats its own discount and posts the same bill as before. And if your workload mentions security tooling, budget for the reroute: you are paying Opus 5.5 prices and may be served Opus 4.8 output.

My take

The story everyone will write this week is the price war. The story I think actually matters is that the frontier labs are now competing on cost per task — not cost per token, which has been falling for two years, but the fully-loaded number that includes how many tokens a model needs to finish the job it started. Opus 5.5 walks right up to that line: it is the smartest model measured this week and it is only cheap when you do not ask too much of it. That is not a knock. It is just the first release where the bill and the leaderboard finally describe different shapes.

What I am watching next: Sonnet 5.5 and Haiku 5.5 landing "in the coming weeks" with the same 40%-style claims, whether anyone ever measures the cyber-fallback rate, and what the next index version does to these numbers. I will post when the family is complete.

Sources: Anthropic — Claude Opus 5.5 announcement · Opus 5.5 system card · Artificial Analysis — Claude Opus 5.5 (and the model pages for Opus 5, Fable 5.1, GPT-6 Astra, GPT-6 Sol) · CodeRabbit — Opus 5.5 evaluation · Hacker News thread (1,394 points; quotes above are from the named commenters) · The Verge · The Decoder · The Implicator · MarkTechPost. What I did not run: every benchmark number in this post is somebody else's measurement — Anthropic's, Artificial Analysis's, CodeRabbit's, or Zapier's — and I have labeled which is which throughout. The arithmetic I did myself was comparing published figures against each other, like the $5.98-vs-$5.86 cost per task at max effort. Earlier on this blog: the GPT-6 Sol launch post, what changed in Fable 5.1, and the Claude Code limit math.

Comments