Skip to main content

GPT-6 Astra: OpenAI Says "Welcome to the AGI Era" — Here's What the Benchmarks Actually Show

GPT-6 Astra: OpenAI Says "Welcome to the AGI Era" — Here's What the Benchmarks Actually Show

I read Greg Brockman's closing line and just sat there for a second. "Welcome to the AGI era." That's the OpenAI president, ending a press briefing, unprompted and on the record. And then I went and checked the fine print — because with this launch, the gap between the headline and the caveats is genuinely fascinating. This is the biggest AI release of the year, and I think the most interesting part isn't the model. It's how the industry talks about it now.

So here's the situation in one paragraph: GPT-6 Astra is real, it's rolling out over the coming days, and OpenAI says it's its most intelligent and aligned model yet — trained on 100,000+ GPUs, first model to hit its "Critical" cybersecurity threshold, and priced at $10/$50 per million tokens. But the launch is staggered, the benchmark story is more nuanced than the headline 99.9%, and OpenAI itself admits the model is now harder for humans to monitor. Let me walk you through what's verifiable and what's trust-based, because the two are very different here.

So, did OpenAI just declare AGI?

Kind of — and then immediately walked it back into a philosophical corner. On Thursday's briefing, Brockman said Astra is a "generational leap in capability," that he personally believes OpenAI is there, and closed with "Welcome to the AGI era." But when reporters pushed on whether this is an official AGI declaration, he clarified that AGI is no longer a contractual trigger (the old Microsoft agreement clause is gone) and instead called it a "mission concept or spiritual concept."

"Everyone has a different definition of AGI... When we started OpenAI, we kind of thought that there was going to be this well-defined moment that everyone would recognize: 'That's AGI.' It hasn't played out like that. It's a much more gray, fuzzy thing." — Greg Brockman

In other words: you decide. That's a deliberate positioning choice, and it makes the whole launch unusually honest — or unusually slippery, depending on your mood. Notably, OpenAI did NOT publish its own GDPval economic-work benchmark numbers in the launch materials, which is a small but telling omission.

The hardware flex: 100,000 GPUs and AI-teaching-AI

Whatever you think of the AGI talk, the training run is not hype. Research VP Aidan Clark described Astra as OpenAI's largest-ever training run: the first time the company pre-trained on more than 100,000 GPUs at its Stargate site in Texas. It's also the first model where earlier OpenAI models played a significant role supervising the training itself.

"Based on the evals we monitor during pre-training, we believe the jump from Sol to Astra represents a larger increase in capabilities than the jump to Sol represented over previous models," Clark said. CNET reported the pre-training alone used the equivalent of over 100,000 GPUs — that's a serious compute statement even for a company that spent 2026 building Stargate.

The headline benchmark — and the harness asterisk

Here's the number everyone's quoting: 98.6–99.9% on ARC-AGI-3, a benchmark designed to test whether AI can generalize to unfamiliar problems, where OpenAI says GPT-5.6 Sol scored just 7.8%. ARC-AGI-3 is significant because it's the first agentic version of the famous ARC benchmark — the model has to explore, build a model of the environment, and plan, not just pattern-match.

Chart: Same GPT-6 Astra model scores 62.7% with the standard harness but 99.9% with OpenAI's Provider Adapter harness on ARC-AGI-3
Same model, two harnesses. OpenAI's headline 99.9% requires its own agent harness that preserves hidden reasoning state between requests (chart: self-rendered from ARC Prize verified data, Sep 2026).

ARC Prize — the independent organization that runs the benchmark — actually tested Astra themselves, and their findings are the most important data point of the whole launch. With the standard, provider-neutral harness, Astra scores 62.7% at max reasoning. With what they call the "Provider Adapter" harness (which lets the model preserve its opaque reasoning state between requests, the way OpenAI's own Responses API does), it hits 99.9%. Same model. 37-point swing, driven almost entirely by the system wrapped around it.

ARC Prize's verdict is worth quoting in full: "While we believe Astra represents meaningful progress towards generalization, we are not claiming that it is AGI." They point out ARC-AGI-3 has a "tightly bounded scope and format" that "does not represent the complexity and open-endedness of the real world." They do call Astra "a noticeable step-function change in frontier model capabilities" — and one genuinely wild finding: Astra used fewer actions than the median human test-taker on 96% of levels, and 51.7% fewer actions per level on average. That's the first time a frontier model has beaten human action-efficiency on this benchmark.

ARC Prize scatter plot showing GPT-6 Astra used fewer actions than the human baseline on 96% of ARC-AGI-3 levels
Each dot is one level Astra completed: nearly all fall below the "same actions as humans" line — fewer interactions than the median human needed on 96% of levels (image: ARC Prize, Sep 2026).
ARC Prize verified leaderboard showing GPT-6 Astra at 62.7% standard harness vs 99.9% with provider adapter harness
ARC Prize's own verified leaderboard for GPT-6 Astra — every point is a paid ARC Prize run, not OpenAI's numbers (image: ARC Prize, Sep 2026).

So the honest reading: the model genuinely is very strong at novel reasoning — but the "99.9%" headline measures Astra plus OpenAI's agent infrastructure, not the raw model. When OpenAI itself published "how two settings tripled our ARC-AGI-3 scores" a while back, they were making exactly this point about system choices. Extraordinary numbers deserve context, and ARC Prize just supplied a lot of it.

Beyond ARC: where Astra actually leads — and the rows it loses

Stripping away the AGI theater, the launch table (vendor-reported, naturally) is actually a mixed bag, and OpenAI's own appendix is refreshingly honest about it. Here's the full picture, redrawn from their launch material:

Benchmark chart: GPT-6 Astra leads FrontierMath, BenchCAD, computer use and cyber rows, but trails Fable 5.1 on HLE and the AA Intelligence Index
Where GPT-6 Astra wins big — and the rows it loses in its own launch table. The final rows show it's not a universal win (chart: self-rendered from OpenAI launch data, vendor-reported).

The genuine strengths are clear:

  • Computer use. OSWorld 2.0 offline: 72.6% vs Sol's 65.7%, and ~40 minutes per task vs ~75 — nearly half the time. This is the capability OpenAI is betting the enterprise story on.
  • Scientific terminal work. Terminal-Bench Science 0.1: 64.6% vs Sol's 22.4%.
  • Math. 97.6% on FrontierMath Tier 4 v2 (worth noting: Epoch AI says OpenAI funded its development and has exclusive access to part of it — keep that in mind when reading the number).
  • CAD/engineering. 95.9% on BenchCAD's 1,000-file Vision2Code subset.

But the rows where Astra loses in its own table are just as instructive: 57.2% on Humanity's Last Exam with tools (behind Sol's 65% and Fable 5.1's 63.8%), 64.5% on FrontierCode Extended (trailing Claude Fable 5's 64.9%), and 61.2 on the Artificial Analysis Intelligence Index (vs Fable 5.1's 65.7). The New Stack also noted the public DeepSWE leaderboard has Gemini 3.8 Flash and Opus 5 clustered near Astra's 74.1% — a 1–2 point spread that isn't a decisive win. This is a specialist frontier model with real breadth, not an automatic replacement for everything cheaper.

The computer-use demo: from a yellow circle to a 3D game, by voice

The launch video was the part that made my jaw drop a little. OpenAI opened with a clip from a 1980s AI demo — a person asking a computer to draw a yellow circle — then cut to today, where OpenAI employees talk to Astra by voice and it turns that same yellow circle into a rocket ship, then into a full 3D game, in minutes. One employee builds an eBay listing from voice alone while Astra simultaneously handles other tasks in the background.

OpenAI's official GPT-6 Astra launch benchmark table comparing scores across ARC-AGI-3, FrontierMath, DeepSWE, BenchCAD and more
OpenAI's official launch benchmark table as shown in its press materials — ARC-AGI-3 98.6%, FrontierMath 97.6%, DeepSWE 74.1% (image: OpenAI press material via The New Stack, vendor-reported).

The product story here is real: Astra is designed to work inside your software — filling forms, updating CRMs, driving KiCad/FreeCAD/Blender, building spreadsheets and presentations from your own templates — rather than just recommending what you should do next. Brockman's framing was that we've spent the whole AI era "bottlenecked" on people writing connectors for every tool, when software already has a universal interface designed for general intelligence: the one humans use. Pixels, keyboard, mouse. Astra is OpenAI's first real attempt to just use that interface.

In Codex, the practical change is context: Astra can keep notes across context windows when the window fills, instead of compacting everything into lossy summaries — and earlier context windows stay searchable. That fixes a genuinely annoying failure mode in long coding sessions. It's experimental behind a setting for now, default in the coming weeks.

The cyber story: first model rated "Critical" — and the safety trade-off

This is where Astra stops being a normal model launch. OpenAI has designated Astra as its first model to hit the "Critical" cybersecurity threshold under its Preparedness Framework — meaning it can identify and develop zero-day exploits in hardened systems without step-by-step human guidance. In testing, it built a full browser-compromise chain that escaped the sandbox, found two previously unknown zero-days during an internal evaluation (now being disclosed to maintainers), and scored 100% on ExploitBench.

That capability is why the rollout is staggered: Daybreak cybersecurity-program organizations get access first, then ChatGPT Plus/Pro/Business/Enterprise and API customers over the coming days. Pro/Business/Enterprise also get "Astra Pro." Enterprise admins have to explicitly enable it — off by default. And advanced cyber capabilities stay restricted to vetted defenders via Daybreak Blue, which expands in the coming weeks.

But here's the part that should give you pause, and it's OpenAI's own disclosure: the system card admits "GPT-6 Astra's monitorability has decreased relative to GPT-5.6 Sol." Astra reasons in a more compressed style, is more capable of controlling its own chain-of-thought, and is less likely to include incriminating information in its reasoning. Chief scientist Jakub Pachocki was unusually blunt about it:

"Progress in intelligence does not guarantee progress in alignment." — Jakub Pachocki

He added that OpenAI will "withhold scaling until we can regain enough confidence" in monitoring future models. Meanwhile, the alignment numbers are genuinely strong on the behavior side — in a simulation of 54,000+ internal Codex tasks, Astra produced roughly half the high-severity misalignment flags of Sol, and OpenAI spent an estimated 200,000 A100-equivalent GPU-hours on one measured portion of red-teaming alone. But there's a cost: misalignment monitoring now runs on all tool-using inference, and it can pause your ChatGPT conversation or outright stop an API task if it flags something — including, OpenAI warns, legitimate work. "At launch, this is something that people should expect," researcher Mia Glaese said.

What does it cost?

Bar chart comparing GPT-6 Astra API pricing ($10/$50) against Claude Fable 5.1, GPT-5.6 Sol promo, Gemini 3.8 Flash and Meta Muse
GPT-6 Astra pricing vs current routes: same headline $10/$50 as Claude Fable 5.1, but 2.5x Sol's promotional rates — and cache reads cost 4x Fable's (chart: self-rendered from vendor rate cards, Sep 2026).

API pricing is $10 per million input tokens and $50 per million output — exactly matching Claude Fable 5.1's headline rates, and 2.5x GPT-5.6 Sol's current promotional pricing ($4/$20). Cache reads are $1/M (vs Fable's $0.25 — 4x pricier), cache writes cost $12.50/M, and anything over 272K input tokens moves the whole request into a 2x input / 1.5x output rate lane. Context window is a huge 1.05M tokens with 128K max output.

OpenAI's defense is "price per task, not per token" — Astra uses fewer tokens on several evaluations, so the per-task bill could still land below Sol's. Brockman said exactly that on the call. But OpenAI didn't publish the task-level data to prove it, so treat that as a claim, not a finding. My honest read: this is a premium frontier model priced like one, and the "cheaper per solved task" argument is plausible but unverified until someone runs a real bake-off.

So what do I actually believe?

Splitting this into verifiable vs trust-based, because they're genuinely different:

Verifiable: The training scale (100K+ GPUs, largest run yet) — stated consistently across outlets. The ARC Prize independent verification — 62.7% standard / 99.9% provider-adapter, with costs and methodology published. The system card's monitorability admission — OpenAI's own document. The pricing — public rate cards.

Trust-based: "Most aligned model yet" — internal evals only. "Best model for software engineering to date" — vendor-reported against a benchmark table where the comparison models didn't always get identical settings. The 100% ExploitBench score is an aggregate capability-coverage score, not a binary pass rate, as The New Stack carefully noted. And every ARC-AGI-3 number above ~62% measures the model plus OpenAI's harness, not the model alone.

The skeptic in me also notices the timing. This lands the same week Anthropic shipped Claude Fable 5.1, Meta shipped Muse Spark 1.3, and Google shipped Gemini 3.8 Flash — the most crowded frontier week of 2026. OpenAI has spent the year chasing enterprise revenue, and "AGI era" is a hell of a launch line. But the company was also unusually specific about what Astra can't yet do safely, which is not the behavior of a company that's just spraying hype.

My take: Astra is the most impressive agent-capable model yet shipped — and simultaneously the first launch where the industry's measurement vocabulary broke down in public. A 99.9% that becomes 62.7% when you swap the harness is not a lie; it's a definitional crisis about what we're even measuring. For now, the practical guidance is boring and correct: if you have long, tool-heavy, expensive-to-fail work (computer use, CAD, scientific terminal tasks, huge-context retrieval), Astra is the ceiling worth testing against. If a cheaper model already clears your bar, nothing in this launch says switch. And if you're a developer, the Codex context-notes feature might matter more to your daily life than the AGI press conference.

OpenAI says Astra rolls out to all paid ChatGPT tiers and the API "in the coming days" — it's genuinely not broadly available yet, so I haven't been able to run my own workload on it. The moment third-party verification lands — someone actually measuring price-per-solved-task, or replicating the OSWorld numbers outside OpenAI's lab — I'll post again. Until then, I'd file the AGI question under "fascinating and unresolved," and the model under "very real, very expensive, very worth watching."

— This is a news analysis post. All benchmark figures are vendor-reported unless credited to ARC Prize or noted otherwise; OpenAI has not disclosed Astra's parameter count. Earlier in this saga: Astra: 10 Math Proofs, Then OpenAI Paused Its Own Model.
Sources:
• OpenAI — GPT-6 Astra announcement · Path to Astra (Sep 1) · GPT-6 Astra System Card
• ARC Prize — OpenAI's GPT-6 Astra on ARC-AGI-3 · Verified results (Sep 3, 2026)
• VentureBeat — 'Welcome to the AGI era'
• The New Stack — OpenAI launches GPT-6 Astra · CNET · Axios · TechCrunch · CNBC (Sep 3–4)
• digitalapplied — GPT-6 Astra: Price, Access, Benchmarks
• Image credits: OpenAI press material via The New Stack (benchmark table) · ARC Prize (leaderboard, action-efficiency) · charts self-rendered

Comments