Skip to main content

Claude Sonnet 5.5: Half the Price of Opus — Except at Max Effort

Claude Sonnet 5.5: Half the Price of Opus — Except at Max Effort

Anthropic's launch card for Claude Sonnet 5.5

Anthropic's launch card for Sonnet 5.5, shipped Monday, September 28. (image: Anthropic)

Anthropic shipped Claude Sonnet 5.5 on Monday, and the price card did not move: still $2 per million input tokens and $10 per million output — exactly half of Opus 5.5's rates. The vendor numbers suggest something that used to cost a generation to buy: Terminal-Bench 4.0 climbs from 10.3% to 70.6%, and on the knowledge-work benchmarks the mid-tier model now finishes within two points of the flagship. The independent board mostly agrees — and it adds the part the launch charts leave out. At medium effort, Sonnet 5.5 beats its predecessor's best score for about an eighth of the cost. At max effort it becomes the worst deal in the family: 56 points at $7.60 per task, against Opus 5.5's 58 at $5.98.

I ended my Opus 5.5 post six days ago with three things I wanted to watch: Sonnet 5.5 landing with the same discount-style claims, whether anyone would ever put a number on the cyber-fallback rate, and what the next index version would do to these numbers. The family is not complete yet — Haiku 5.5 is still "in the coming weeks" — but two of the three just landed, and one of them landed inside a PDF footnote where nobody would look. That story is below, along with what the model is actually like to use, and where the "30% cheaper" claim holds and stops holding.

What actually shipped

Sonnet 5.5 is closed weights, model ID claude-sonnet-5-5, available on Anthropic's own platform plus Amazon Web Services, Google Cloud and Microsoft Azure, with zero data retention available. It keeps the 1M-token context window, and it is the second model in the "Claude 5.5 family" after Opus 5.5. The reliable knowledge cutoff is June 2026.

Price per 1M tokensSonnet 5.5Sonnet 5 (replaced)Opus 5.5 (sibling)GPT-6 Sol (same-week rival)
Input$2.00$2.00$4.00$2.00
Output$10.00$10.00$20.00$10.00
Cache reads$0.20$0.20$0.20$0.20
Cache writes$2.50$2.50$5.00$2.50

Anthropic's pitch is not a cheaper sticker — the sticker is identical to Sonnet 5's. It is that the model needs fewer tokens and fewer tool calls to finish the same job, "up to 30% less per task" in their testing, while generating output 30%+ faster. That framing will sound familiar if you read my Opus 5.5 post: the industry has quietly moved from competing on cost per token to cost per task, and this release is the first one where the mid-tier gets the same argument.

The effort dial is now the whole interface. Five levels — low, medium, high, xhigh, max — with Medium as the default in Claude Code and the Claude apps, and High as the default on the Claude Platform. Adaptive thinking is always on; the old "turn thinking off" switch now returns a 400 error, and what remains is a between_tools setting that keeps only the up-front thinking off. Forced tool calls are retired, the minimum cacheable prompt drops to 512 tokens, and non-default temperature or top_p values are rejected outright. If you maintain an integration against Sonnet 5, the migration guide is not optional reading.

A screenshot from Anthropic's launch page showing the previous model, Sonnet 5, mid-write on a starling murmuration program

Anthropic's launch page has the models write their own demos. This is the older Sonnet 5 mid-write on a starling murmuration program — 4,520 tokens in, still building the boids loop. (image: Anthropic)

The starling murmuration render produced by Sonnet 5.5, with an overlay reading Output 4,158 tokens, Flying 12.5 s

And this is the 5.5 render of the same prompt, with the launch page's own overlay: the program written and running, 4,158 tokens of output, 12.5 seconds of flight. It also demonstrates something benchmarks do not — Anthropic says this is the first Sonnet that beats Pokémon Red from screenshots alone, which is either a great milestone or a great meme, and honestly it is both. (image: Anthropic)

The numbers Anthropic sent out — and the asterisk on the crown

Here is the vendor table as shipped. The rows that matter for the mid-tier story are Terminal-Bench, CursorBench and the two knowledge-work evals, because those are the ones where Sonnet 5.5 either closes on Opus 5.5 or beats it outright — both are vendor-reported, so read them as claims rather than measurements.

BenchmarkSonnet 5.5Sonnet 5Opus 5.5GPT-6 Sol
Terminal-Bench 4.070.6%10.3%66.4%*not reported
FrontierCode 1.1 (Main)52.1%*42.4%54.4%49.3%
CursorBench 4.055.5%34.1%57.8%not reported
GDPval-AA v2.11844144918461487
AA-Briefcase v1.11811135918221483
OSWorld 2.1 (partial)80.1%57.0%81.8%not reported
Chartography (no tools)61.6%15.6%64.4%53.6%

* Sonnet 5.5's FrontierCode row is its xhigh score; at max effort it actually scored lower (46.2), because it kept launching Claude Code's multi-subagent code-review skill and blew past the task scope — Anthropic's own footnote, and a nice preview of where this dial stops making sense. Opus 5.5's Terminal-Bench row is its xhigh result.

That Terminal-Bench row is the one everyone reposted, and it deserves the footnote treatment. The system card spells out the conditions: 66 tasks, five trials each, and both models run with safeguards live, which means some trials were answered by a fallback model instead of the model being tested — 1.5% of Sonnet 5.5's trials, against 10% of Opus 5.5's. Add the standard errors (±2.5 and ±2.6 points) and the six-point gap that reads as "Sonnet beats Opus at coding" is really "the two models land within noise of each other, and Opus was handicapped more by its own rerouting". Anthropic is transparent about all of this — the numbers are printed, not buried — but the headline is not the one the table produces.

Two-panel bar chart of vendor-reported benchmark scores for Sonnet 5.5, Sonnet 5, Opus 5.5 and GPT-6 Sol

The vendor table as a picture, with the blanks kept blank. Note the row where the mid-tier model genuinely wins: AutomationBench, where Sonnet 5.5 scores 44.7 against Opus 5.5's 42.5 — and the row where it is close but clearly behind: CursorBench, measured independently by Cursor. (chart: Open Source Factory; vendor-reported figures)

Two more vendor findings worth keeping: with tools, Sonnet 5.5 actually edges Opus 5.5 on chart recognition (90.2 to 89.0 — the no-tools gap goes the other way, which tells you tool use, not eyesight, is where the mid-tier caught up). And the earnings-deck test in the announcement — quarterly filings plus call transcripts plus a slide template, first draft judged send-ready by two reviewers — is the kind of claim benchmark tables cannot make, so make of it what you will.

What independent measurement says

Artificial Analysis runs every model at every effort level in one harness, which makes its board the closest thing to an apples-to-apples ruler for a release whose entire pitch is a dial. Sonnet 5.5's ladder, as measured there:

Effort settingIntelligence IndexCost per taskTokens on the run
low36$0.41—
medium (default in Claude Code / apps)41$0.5929M
high (default on the platform)47$1.0850M
xhigh52$2.74100M
max56$7.60410M

The claim arithmetic checks out at the setting most people will use. Sonnet 5's best score on the same index is 38, at $5.09 per task. Sonnet 5.5 at medium scores 41 — three points better — for $0.59, about an eighth of the cost. That is as close to a free upgrade as this market produces: the thing you are told to run by default is both smarter and roughly 8x cheaper per completed task than the thing it replaces. (The low setting, 36, does not quite clear Sonnet 5's best — the upgrade lives at medium and up, not at the floor.)

Scatter chart of Artificial Analysis Intelligence Index versus cost per task, showing the Sonnet 5.5 and Opus 5.5 effort ladders plus Sonnet 5, GPT-6 Sol and Fable 5.1

Color inside the lines: the orange curve is Sonnet 5.5 at its five effort levels, the dashed blue is Opus 5.5 at its five, and the standalones (Sonnet 5, GPT-6 Sol, Fable 5.1) sit at their max-effort runs. Opus 5.5 is above Sonnet 5.5 at every shared cost — until the last point, where the ordering flips in the way the annotation describes. (chart: Open Source Factory, from Artificial Analysis)

Now the part the launch charts skip. Compare Sonnet 5.5's ladder against the sibling it undercuts:

  • At xhigh, Sonnet 5.5 scores 52 at $2.74 per task. Opus 5.5 at high scores 54 at $1.82 — two points more, roughly a third less. The comment thread found this within hours.
  • At max, Sonnet 5.5 scores 56 at $7.60 while Opus 5.5 scores 58 at $5.98, and burns 410M tokens to Opus's 260M — 58% more. At the top of the dial, the mid-tier model is the more expensive way to get a lower score.

None of this contradicts Anthropic's claims, because Anthropic's claims are about the default settings — and at default settings the value is real. It does mean the family sorts differently than the naming suggests: below ~$1 per task, Sonnet 5.5 is the best deal Anthropic has shipped this year; above that, you are paying mid-tier prices for what is increasingly an Opus-shaped workload, and the honest move is to switch models rather than keep cranking. And on the same-harness comparison against the model it shipped alongside: Sonnet 5.5 at high (47 at $1.08) essentially ties GPT-6 Sol at max (48 at $1.06) — a dead heat at a dollar a task, which is a way of saying the $2/$10 tier is now genuinely contested.

The efficiency claim, metered

The independent verbosity receipt complicates the story in the same direction the ladder does. Artificial Analysis counts the output tokens each model spends to finish its full index run, and Sonnet 5.5 at max effort burned 410M — more than Sonnet 5's 370M, and far more than Opus 5.5's 260M. "Fewer tokens per task" is true at the defaults; at the ceiling, the new model is the hungriest thing in the family.

Four-panel chart: tokens per Intelligence Index run at max effort, Base44 iterations per build, Balyasny tokens per answer, and tester-reported gains

Four receipts. Left: the independent token meter at max effort, with each model's cost per task in the second line — the cheapest run here is GPT-6 Sol's, not the Sonnet's. The other three panels are the hoard of efficiency numbers Anthropic collected from customers, and they all point the same way: at working settings, the new model does the job in fewer steps. (chart: Open Source Factory, from Artificial Analysis and vendor-supplied tester figures)

The tester numbers are the most convincing part of the announcement, precisely because they are unglamorous: Box reports 2.4x faster results with 12% fewer total tokens; Slack, 14% fewer output tokens on its internal evals with no prompt changes; Zendesk, tickets processed 20% faster; Lovable, about a third fewer tool calls and half the shell executions; Balyasny, 121,000 tokens per finance answer instead of Sonnet 5's 497,000. Base44's number is the one I keep thinking about: across 118 real app builds, 3.6 iterations per build against 7.7 for Opus 5. "Iterations per build" is not a benchmark anyone optimizes for, which is exactly why it is informative.

Simon Willison ran his usual pelican test and produced the cleanest picture of the dial that exists. The same prompt at each setting: low and medium finished in ~10 seconds with zero thinking tokens; high spent 745 thinking tokens, xhigh 2,535; and max burned through 128,000 thinking tokens over 15 minutes and 40 seconds — and ran out before drawing the pelican, for $1.28. The effort dial is not a smooth curve. It is a flat line with a cliff at the end, and the cliff is where the pricing math inverts.

First Sonnet with Opus-grade brakes — and the number I asked for

Sonnet 5.5's cyber capabilities improved enough that Anthropic gave it the flagship treatment: it is the first Sonnet model to launch with the cyber safeguards and fallbacks built for Opus-class systems. In practice, higher-risk cybersecurity requests get quietly answered by Sonnet 5 instead — automatic in Anthropic's own products, opt-in on the API — while routine bug finding and fixing is untouched. Three classifier stages gate it (an activation probe, an on-model classifier, and a trained LLM classifier). On Anthropic's cyber-harm coverage set those classifiers keep 99.43% recall, against 82.63% for Sonnet 5. Robustness improved too: on the company's internal rewind-attacker evaluation the attack success rate fell from 57.2% to 21.0% — though that figure is still far worse than Opus 5.5's 4.0%, a trade Anthropic says it made deliberately, since less-robust safeguards are acceptable on a model that is less cyber-capable to begin with. It is also the first Sonnet with classifiers against reasoning-extraction distillation, and preserved thinking now ties a model's thinking to the account that created it.

But here is the thing I promised to watch for, and it is now measured. The system card puts concrete numbers on the fallback routing for the first time: on Terminal-Bench 4.0, 1.2% of Sonnet 5.5's requests were answered by a fallback model, affecting 1.5% of trials — and Opus 5.5's otherwise-stronger table position was dragged by a 2.5%/10% rate of its own. That is one benchmark, not production telemetry, and nobody has published what the routing rate looks like on real traffic. Still, after a year of "we reroute some requests" with no denominator, a published denominator — even a small one — is a real change, and it retroactively justifies the skepticism about the earlier Terminal-Bench comparisons.

"Paying $200 a month and part of their Cyber Verification Program but can't use Opus 5.5 or Sonnet 5.5 for any authorized bounty work. Immediately get flagged for `Cyber`. This is bollocks." johnmlussier, on the Sonnet 5.5 thread (Hacker News, thread 49881850)

That complaint is the honest tension in the safeguards: a paying customer with credentials for the approved-defender program still hits the classifier. Anthropic's answer is the expansion of that program — tiered access for vetted defenders on Sonnet 5.5, Opus 5.5 and the Mythos models "in the near future" — which is to say the routing gets less annoying, but only for people who apply. One detail worth appreciating in the other direction: over-refusal on benign-but-sensitive requests collapsed from 0.59% to 0.02% on the API, so the model is simultaneously more guarded where it matters and less jumpy where it does not.

What the room actually said

The launch thread was warmer than the last few Claude releases, with the same two complaints repeating. The warmth, first, because it is verifiable: early users found it fast and competent, and the cost-per-score shape of the independent charts is genuinely new for a Sonnet.

"Big jump on Agentic coding from 10.3% -> 70.6% ... Opus 5.5 is really strong so this is impressive especially for the cost." pavitheran, Sonnet 5.5 thread (Hacker News)

The skeptical takes were sharper, and mostly about the middle of the dial — the exact region the launch charts draw through. One commenter asked the question I would frame the whole release around: "Why would you use Sonnet 5.5 on xhigh if you would get better results (higher score, cheaper cost) on Opus 5.5 high?" Another put it more bluntly — once you are at high or xhigh, you are better off on Opus at low or medium. And one noted that in almost every configuration on the cost-performance chart, the new Sonnet "looks worse than Opus" — which is true in the sense that matters (dollars per point) and false in the sense the launch is selling (dollars per task at the default). Both readings are of the same chart. That is what a genuinely contested product looks like.

Two smaller threads worth pulling: one commenter observed that Fable is missing from every vendor chart ("Maybe Fable is out the door?"), and another pointed out that MiMo V2.6 Pro — an open-weights model — basically matches Sonnet 5.5 at high effort for far less. On the first, I have nothing but the observation; on the second, that is the price tier doing what it was always going to do. The $2/$10 band now contains GPT-6 Sol, Gemini 3.8 Flash (introductory $0.75/$3.75 through the end of the year), and a new Sonnet that is Opus-adjacent on knowledge work. Mid-tier is where the war is.

So should you switch?

If you use Claude apps or Claude Code, the upgrade is nearly free: same price, better model, and the default Medium setting is exactly where the value math works. The days of reaching for Opus on well-scoped work are mostly over — that is the whole point of this release, and unlike the vendor's benchmark rows, the cost-per-task version of the claim survives an independent check.

If you pay API rates, run one number before believing any discount: your tokens-per-finished-task ratio at your own effort setting. Teams at the defaults — especially cache-heavy agent loops, where the $0.20 cache-read price is the real bill — are looking at meaningfully cheaper work, and the customer reports above suggest the tool-call reduction is real. Teams that crank effort to max for hard tasks should look at the ladder instead: at that end, Opus 5.5 is both the better model and the cheaper one, and the correct optimization is to switch models, not settings. And if your workload trips the cyber classifiers, budget for the reroute to Sonnet 5 — the fallback is transparent at the request level, but it happens more often than the launch week discussion assumed.

My take

Six days ago I wrote that the interesting thing about Opus 5.5 was that its bill and its leaderboard described different shapes. Sonnet 5.5 is the release where those shapes finish diverging, and it is a better product for it. The mid-tier is now Opus-adjacent on knowledge work, roughly a tenth of the cost of what it replaces at the settings people actually use, and meaningfully cheaper to finish real jobs with — and it still has a ceiling where the flagship wins on both axes, cleanly. That is what a sane model family is supposed to look like: pick the workload first, then pick the model, then pick the dial. Anthropic spent a year selling the dial; this is the first release where the dial has an obvious wrong end.

What I am watching next: Haiku 5.5 when it lands, Artificial Analysis's speed measurement for this model (still unmeasured — which is itself a gap in a launch pitched on "30%+ faster"), and whether anyone outside Anthropic publishes what the cyber-fallback rate looks like on production traffic. I will post when the family is complete.

Sources: Anthropic — Introducing Claude Sonnet 5.5 · Sonnet 5.5 system card (PDF) · Claude Platform docs — Sonnet 5.5 · Artificial Analysis — Claude Sonnet 5.5 (and the model pages for Opus 5.5, Sonnet 5, GPT-6 Sol, Fable 5.1) · OpenRouter · Hacker News thread 49881850 (quotes above are from the named commenters) · TechCrunch · VentureBeat · Unite.AI. What I did not run: every benchmark number in this post is somebody else's measurement — Anthropic's, Artificial Analysis's, or Cursor's — and I have labeled which is which throughout. The arithmetic I did myself was comparing published figures against each other, like the $7.60-vs-$5.98 cost per task at max effort and the eighth-of-the-cost ratio at medium. Earlier on this blog: the Opus 5.5 launch post, Opus 5.5 vs GPT-6 Astra on tokens and bills, and the GPT-6 Sol launch.

Comments