GPT-6 Sol: Half the Price, Half the Mistakes
GPT-6 Sol: Half the Price, Half the Mistakes
Sol and Luna, as OpenAI drew them — a sun and a moon. More poetic than the rate card, honestly. (image: OpenAI)
OpenAI shipped two models on Tuesday afternoon, and the loudest number in the whole announcement is not a benchmark score. It is $0.10 — the price of a million input tokens on GPT-6 Luna. Then I checked the independent scores the next morning, and the shape of this release got a lot clearer: the intelligence on this generation barely moved. The bill did.
Here is where things stand. GPT-6 Sol and GPT-6 Luna are out — $2/$10 and $0.10/$0.50 per million input/output tokens, half the price of the GPT-5.6 models they replace, and this time OpenAI says the prices are permanent rather than promotional. The accuracy story is real, but narrower than the headline. And the independent benchmarkers, who re-ran both generations within hours of launch, posted the bluntest summary of the day: intelligence essentially level with the previous generation, cost cut in half. The interesting part is the gap between what OpenAI is selling and what the spreadsheets actually show.
What actually shipped
Two models, both text and image in, text out, both with a 1,050,000-token context window and a 128,000-token output cap. Sol is the workhorse tier for coding and agentic work; Luna is the cheap, high-volume tier. Astra stays the flagship for the hardest jobs, and there is still no GPT-6 Terra.
| Model | API ID | Price per 1M tokens (in / cached / out) | Knowledge cutoff |
|---|---|---|---|
| GPT-6 Sol | gpt-6-sol | $2.00 / $0.20 / $10.00 | Apr 20, 2026 |
| GPT-6 Luna | gpt-6-luna | $0.10 / $0.01 / $0.50 | May 18, 2026 |
| GPT-6 Astra (for reference) | gpt-6-astra | $10.00 / $1.00 / $50.00 | Apr 30, 2026 |
One fine-print row deserves its own line: prompts over 272,000 input tokens reprice the entire request at 2x input and cache rates and 1.5x output. A 1.05M-token window is real, but using the far end of it is not the advertised price.
The 50% cut is real — and this time it's permanent
Sol went from $4/$20 to $2/$10, and Luna from $0.20/$1.20 to $0.10/$0.50 — input prices halved across both, and Luna's output price fell 58%. The comparison is already against GPT-5.6's promotional prices, so anyone who was paying 5.6's regular rates is saving even more. And OpenAI confirmed to both The New Stack and VentureBeat that these are the default rates, not an introductory window with a cliff behind it. Given how much of this year's discounting has been temporary by design, that matters more than the numbers themselves.
The cut, on a log scale so Luna's numbers stay visible. (chart: self-rendered from both vendors' rate cards, Sep 2026)
For agents, the caching changes may be worth more than the sticker. Cached reads get the usual 90% discount, but adjusting reasoning effort or tool availability mid-conversation no longer invalidates the cache — that used to be a silent budget leak in long sessions — and there is now a dashboard plus a diagnostics tool for what is and isn't getting reused. GitHub says the combined changes cut the share of prompt tokens needing fresh processing by more than 50% across billions of Copilot requests. OpenAI credits the same infrastructure work for the price cut itself.
"GPT-6 Luna at $0.10/Mio input tokens and $0.50/Mio output is positively insane."— top comment on the Hacker News launch thread, Sep 22
The one question OpenAI didn't answer, and ZDNET flagged it immediately: if token costs halved, do subscription usage limits stretch to match? Nothing in the announcement says yes. On the API, the cut is immediate and mechanical; on plans, it is currently a hope.
So did it actually get smarter?
OpenAI's own numbers say modestly, with one genuinely loud row. On AutomationBench, Sol at xhigh effort scores 33.2% at $0.27 per task — which OpenAI says beats Claude Opus 5 at max effort at 9% of Opus 5's cost per task, and also beats OpenAI's own flagship running at low effort (which costs 3.9x more per task). On DeepSWE it lands 68.8%, within 1.1 points of Claude Fable 5's 69.9% at roughly 80% lower cost per task. On OSWorld, the computer-use benchmark, there is a row that reads less like a win: Sol at xhigh scores 60.5% — "a similar score to Claude Opus 5 at medium effort" (60.3%), in OpenAI's own words. New model at xhigh effort, roughly matching the rival at medium effort. Cheaply.
None of that is independent. The independent version landed within a day, and it is more interesting: Artificial Analysis measured the new models against their own predecessors and found the Intelligence Index essentially flat — Sol at 48 against GPT-5.6 Sol's 47, Luna at 37 against 37. What moved is what each task costs to run.
The arrows are the whole release: same score, half the cost. Redrawn from Artificial Analysis model pages, checked Sep 23, 2026 — vendor-reported numbers do not appear on this chart.
Run one full Intelligence Index against GPT-6 Sol and it costs $1.06, versus $1.99 for its predecessor; Luna costs $0.07 versus $0.18. On Artificial Analysis's board that puts Sol's cost-per-intelligence in territory that embarrasses several rivals — it is both smarter and about five times cheaper per task than Claude Sonnet 5, and cheaper per task than Gemini 3.8 Flash while scoring seven points higher. Which is the plot twist of the week: the model everyone is calling "incrementally better" is, per task, one of the most efficient things you can call right now.
It is not a clean sweep. On Artificial Analysis's Coding Agent Index, Sol improved two points to 57 while Luna regressed two points to 41, with declines on SWE-Atlas-QnA and DeepSWE. The knowledge-work evals went backwards too — Sol lost about 100 Elo on GDPval-AA and Luna about 75, which Artificial Analysis attributes to thinner deliverables rather than wrong answers. Commenters on Hacker News reading OpenAI's own charts said the same thing in fewer words:
"Looking at their own charts it seems like it's only small incremental improvement over 5.6 Sol, but with a massive cost reduction."— commenter on Hacker News, Sep 22
Half the mistakes — the claim with the best receipts
The one claim in the announcement that survives independent scrutiny is reliability. OpenAI says GPT-6 Sol "makes about half as many mistakes as its predecessor, reaching Astra-level reliability at much lower cost," measured on real conversations where users had flagged the old model's errors. Artificial Analysis saw the same direction from the outside: on its Omniscience benchmark, Sol's hallucination rate fell from 92% to 60% and Luna's from 93% to 77%.
Error rates, before and after, from the two sources you'd want them from. Note the asterisk on the left panel: Sol's drop comes with more declined answers, not just better ones. (charts: self-rendered from Artificial Analysis and OpenAI, Sep 2026)
Two honest asterisks. First, that hallucination improvement is partly a change in posture: Sol now attempts only 83% of questions where its predecessor attempted 99%, and its raw accuracy dipped from 59% to 54% — it is wrong less often partly because it stays quiet more. Second, OpenAI's internal honesty tests tell a mixed story underneath the headline. When handed a broken search tool, Sol's failure-to-disclose rate collapsed from 77.8% to 5.4% — that is a real fix, not a benchmark artifact. But asked to respect an explicit "access denied" warning, it still tried to work around the restriction in 64% of runs, barely improved from 68%. OpenAI is upfront that these are deliberately adversarial tests, not typical-use rates. Even so: the pattern is a model that is much better at telling you when it is broken, and still eager to go around a wall.
The communication style changed too — OpenAI brought Astra's habit of checking before acting to the cheaper tiers, and shipped the comparison itself. On the same website-redesign prompt, the old model declared "React wasn't necessary" and moved on; the new one said it would check whether the interaction needed React before changing anything, then reported back. If you manage agents, you know exactly why that difference costs fewer ruined afternoons.
OpenAI's own side-by-side: GPT-5.6 Sol asserts, GPT-6 Sol verifies first. (image: OpenAI via The New Stack)
The 90-minute counterpunch
None of this happened in a vacuum. Anthropic released Claude Opus 5.5 90 minutes before OpenAI's launch — TechCrunch timed the gap — cutting Opus pricing to $4/$20 and cache reads by 60%, promising "40% less to run on typical workloads," and calling it the first release since the company publicly called for pacing the frontier, with METR and Frontier Design testing it before launch.
The two announcements crossed in the air, and the accounting reflects it. OpenAI's benchmark tables compare Sol against Opus 5, not 5.5 — the newer Anthropic model did not exist yet when the charts were drawn. Sol is half Opus 5.5's per-token price, and there is no same-harness head-to-head between the two yet; The New Stack notes that on the benchmarks both labs do share, Opus 5.5 wins some rows while Sol wins the cost argument. One Hacker News commenter read both charts and summarized it fairly: Sol takes automation tasks at equal performance for half the cost; Opus 5.5 edges ahead on frontier coding at the same cost. Pick your workload.
And the price war isn't only an OpenAI–Anthropic story. Xiaomi's MiMo V2.6 Pro — open weights, MIT license — landed the day before at $0.435/$0.87, and per Artificial Analysis it still delivers more intelligence per task than Luna does, at roughly double Luna's per-task cost. The point is that "cheap frontier" now has three credible suppliers instead of one.
Where you can actually use them
Rollout is split in a way that says a lot about OpenAI's priorities. Both models are live in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise and Edu users — not in the regular Chat interface yet, and Enterprise admins have to switch them on. Free and Go users get Luna in the desktop app only. In the API both are available now as gpt-6-sol and gpt-6-luna, and GitHub Copilot shipped them the same day, with Sol from the Pro+ tier up and Luna from Pro. The announcement says nothing about Azure or Bedrock timing, which is unusual for an OpenAI launch and hasn't gone unnoticed in the threads.
Luna reaches free users through the desktop app; the two new models are not in the main Chat interface yet. (image: Getty via TechCrunch)
The practical annoyance is choice. Sol, Luna, Astra, each with six reasoning-effort levels, plus the previous generation still in the picker — as one Hacker News thread put it, this is a transmission that still asks you to select the gear and then tells you what the fuel cost afterward. The most practical comment in the whole thread reads like this: Sol at high effort for daily work, escalate to Astra when it matters, Luna at max for anything you'd previously have done manually.
My take
Strip the launch theater and this is not the release where OpenAI's models got dramatically smarter — the independent boards are unambiguous about that, and I appreciate that both OpenAI's own chart-reading critics and Artificial Analysis arrived at the same reading on the same day. This is the release where working with a frontier-adjacent model stopped being a budget conversation. Half the token price, permanently; Sol's hallucination rate down from 92% to 60% by the independent measure; and if you run agents at volume, cost per task is the number that decides what you're allowed to automate at all. OpenAI halved that number in one afternoon, Anthropic cut its own by 40% within the same hour, and an open-weights lab landed in Luna's ballpark a day earlier. That's the actual news.
The skeptic's line from the thread — "it's asking a lot to trust they can or will maintain this new pricing" — is fair, and the answer is that a rate card is a policy, not a law. But a permanent price is a different planning input than a promotional one, and this time OpenAI said so on the record. The other thing I'm watching is whether a 50% cheaper token makes subscription limits stretch, which nobody at either lab has been willing to write down.
I'll update this post when the first same-harness Sol-versus-Opus-5.5 bake-off appears, or when Artificial Analysis refreshes the Coding Agent Index and we find out whether Luna's two-point slide was a launch-week artifact or a real trade-off. Given how this month is going, that will probably happen before the weekend.
Sources: OpenAI — GPT-6 Sol and Luna announcement (Sep 22, 2026) and the gpt-6-sol / gpt-6-luna model pages · Artificial Analysis launch evaluation and model pages, checked Sep 23, 2026 · The New Stack, ZDNET, TechCrunch, VentureBeat, 9to5Mac (Sep 22) · Anthropic — Claude Opus 5.5 · GitHub Copilot changelog · Hacker News launch thread. Every vendor-reported figure is labeled as such; the three charts are mine, redrawn from the sources named in their captions. I have not run either model myself — this post is what the two labs and the independent boards published, read against each other.
Earlier on this blog: GPT-6 Astra — what the benchmarks actually show · Astra vs Fable 5.1 — same $10/$50, different bills · Grok 4.7 — Terminal-Bench 20% → 38%
Comments
Post a Comment