Astra vs Fable 5.1: Same $10/$50, Different Bills
Astra vs Fable 5.1: Same $10/$50, Different Bills
Two frontier models shipped 48 hours apart with the exact same price tag, and neither company benchmarked against the other. OpenAI's GPT-6 Astra landed September 3 calling itself the AGI era; Anthropic's Claude Fable 5.1 landed September 1 with a quiet cache-price cut doing the loud work. I have covered both launches separately on this blog, and the comparison is the post I kept getting asked for, so here it is, with both sides getting their real wins.
The one-paragraph version: Astra takes most of the shared benchmark rows, usually by real but thin margins, and owns computer use, math, and cyber outright. Fable 5.1 takes broad reasoning decisively, holds both independent Artificial Analysis crowns, and wins the rate card on big requests. Then the measured cost per task flips the price story upside down. My read after laying every shared row side by side: Astra takes this one on points, and the scenarios at the end spell out where that verdict holds and where it doesn't.
Where Astra wins: six shared rows plus its own turf
Start with the rows both vendors published or agree on, because those are the only ones that count as a fight rather than a press release.
Astra's genuine territory is wider than this chart. FrontierMath Tier 4 at 97.6% against 87.8% is a real math gap, though it carries the Epoch funding caveat I flagged in the Astra post: OpenAI funded the benchmark's development and holds exclusive access to part of it, so keep that in mind while reading the number. Computer use is Astra's dimension almost by default: 72.6% on the OSWorld offline set with no Fable 5.1 number published against it, and 92.7% ScreenSpot-Pro grounding. Anthropic's own 77.9% partial comes from different grading and cannot be set beside it, which I say explicitly so nobody builds a chart out of it. And cyber is not a contest: 100% ExploitBench aggregate coverage plus 86 of 226 FrontierCyber challenges solved independently, against Fable 5.1's deliberate find-but-don't-exploit posture. Different philosophies, and Astra's ships behind Daybreak gates for exactly that reason.
Where Fable 5.1 wins: the reasoning crown, twice verified
Fable's side of the ledger is narrower in row count and heavier in independence, which is the more interesting combination.
Humanity's Last Exam with tools is the cleanest Fable win in the whole matchup: 65.0% against 57.2%, roughly eight points, and it survives every sourcing variant (OpenAI's own table had it 63.8%, DataCamp 65.0%; either way Fable clears Astra). Artificial Analysis then backs the direction independently: Fable 5.1 at 66 on the Intelligence Index, the highest score it has ever measured, against Astra's 61. One honest footnote rides along: Anthropic's default server-side fallback routed roughly 4% of output tokens through Opus 4.8 or Opus 5, so that 66 is Fable 5.1 as you actually call it, not the raw model in a vacuum. And on the coding-agent side, Fable 5.1 in Claude Code at 70 is the highest Coding Agent Index score on record, against Astra's 67 in Codex, with the harness caveat doing real work: different scaffolding owns part of that 3-point gap, and Astra gets there spending far fewer tokens.
Same sticker, different bill
This is the section that surprised me most while assembling the numbers, and it is why the title leads with price instead of benchmarks.
Read the left panel first. Input, output, cache writes, and batch discounts match to the cent; the only rate-card daylight is cache reads, $1.00 against $0.25, and the long-context surcharge above 272K input tokens where Astra doubles to $20/$2/$75 for the whole request while Anthropic charges standard rates across the full 1M window. On identical tokens that means small requests tie exactly ($22.50 both on the balanced workload), big requests punish Astra (+83% on over-threshold retrieval), and cache-heavy loops favor Fable by $75 on the modeled shape. That threshold detail deserves emphasis: 272K is roughly a quarter of Astra's own 1.05M window, so anything using more than a quarter of the context it ships with bills at the higher tier.
Then read the right panel, because it flips the story. Measured per Intelligence Index task, Astra costs $1.67 against Fable's $3.76 at max effort: 44%, less than half, at the same list price. The mechanism is token frugality. Astra spends far fewer tokens per solved task (a third of Sol's in Codex at max effort, a fifth of Opus 5's at xhigh), while Fable 5.1 runs ~1.7x the output tokens of its predecessor, which is why it costs 20% more per task than Fable 5 despite the cache cut doing real work (without the cut it would be roughly $5.16). OpenAI's own task-cost captions point the same way, for what a vendor self-report is worth: ~31% below on Terminal-Bench Science, ~63% on Terminal-Bench 4.0, ~86% on BenchCAD, at effort settings OpenAI chose. So the rate card favors Fable wherever cache or context dominates, and the meter favors Astra wherever output volume dominates. Which half matters is a property of your workload, not of the models.
The parts neither table shows
Three things that changed my own mental ranking while writing this. First, the harness asterisk applies to both champions, not just Astra's ARC number. The 62.7%-vs-99.9% ARC Provider Adapter swing is the famous one, but the AA Coding Agent gap (70 vs 67) also mixes Codex against Claude Code scaffolding, and OpenAI admits Astra reasons in fewer, more compressed steps that are harder to monitor. System-inclusive scoring is the water both fish swim in now. Second, availability is genuinely asymmetric: Fable 5.1 was GA on day one with an effort dial (low/medium matching Fable 5 quality cheaper) while Astra rolled out staggered behind Daybreak gates with enterprise admins having to opt in. If you need the model this week rather than eventually, that is not a footnote. Third, migration has teeth on the Anthropic side: forced tool_choice now 400s, thinking blocks unreadable to older models, append-only history. Budget a migration afternoon, not just a string swap.
Verdict: Astra takes it on points
I'll say it plainly: on the numbers in this post, Astra looks like the stronger model right now. Seven of eight shared rows, the computer-use dimension largely to itself, the cyber bracket, and less than half the measured cost per solved task. Fable 5.1's wins are real, HLE by eight points and both independent crowns, but they read as the narrower specialist case next to Astra's broader sheet. That is a decision on points, not a knockout, and a matched-harness replication round could still move it.
The default pick is Astra. Agents that click through real software, math and science work, CAD and artifact output, defensive-security tooling behind Daybreak, plus the lowest measured dollars per solved task. Carrying the caveats along: harder-to-monitor reasoning by OpenAI's own admission, the 272K surcharge biting giant requests, staggered availability.
Fable 5.1 is the pick in three shapes. Long cache-dominated loops where quarter-price cache reads decide the bill, giant retrieval requests that live above 272K at standard rates, or Claude Code shops where Fable 5.1 holds the top native-harness score and GA with cost-first effort levels. Its caveats travel with it: 2.25x the measured per-task cost where output dominates, migration breaking changes, fallback-tinted index score at the edges.
Everyone else checks the meter first. If a cheaper model already clears your bar, both launches agree you should stay put: Astra's own HLE loss and Fable's own output-volume bill say the frontier premium needs a workload reason. Run your traffic shape through both panels of the price chart before committing; identical $10/$50 stickers do not survive contact with measurement, in either direction.
My take: the crowded-frontier week of September 2026 gave us the first matchup where the vendor tables and the independent tables disagree in opposite directions at once, and the honest read is that Astra leads on breadth and cost per task while Fable 5.1 leads on reasoning depth and giant-request pricing. I will re-run every number here the moment independent coding-agent replications land in matched harnesses, because the 70-vs-67 with mixed scaffolding is the single row I trust least. Until then, Astra on points.
Sources: OpenAI Astra launch + system card + API rate card; Anthropic Fable 5.1 announcement + platform docs; ARC Prize verified Astra runs; Artificial Analysis Intelligence/Coding Agent indices and per-task costs; DataCamp comparison (rate-card arithmetic, workload shapes, hands-on test); The New Stack (launch tables, settings notes). Earlier on this blog: GPT-6 Astra launch post · Claude Fable 5.1 post · $10 coding subscriptions. Images: OpenAI via The New Stack; Anthropic; charts self-rendered (2 new + 3 reused). Prices checked Sept 5, 2026.
Comments
Post a Comment