OpenAI shipped GPT-6 Astra on September 3 with a one-line pitch: "Anything you can do on a computer, Astra can do for you. Fast."

Greg Brockman went further. "I think it's not unreasonable to feel that we are now in the AGI era." Tobi Lutke replied to the benchmark chart with a single word: "Singularity."

Then the numbers arrived, and the most important one is not on OpenAI's chart. On ARC-AGI-3, Astra scored 99.9%. It also scored 62.7%. Same model, same benchmark, same day. The only thing that changed was the harness it ran inside.

Fig. 1
One model, one benchmark, two harnesses
The harness moved the score 37 points. ARC-AGI-3 SEMI-PRIVATE SAME MODEL. SAME TASKS. DIFFERENT SCAFFOLD. STANDARD HARNESS 62.7% $26,098 to run Transparent notes the model keeps Provider-neutral, comparable PROVIDER ADAPTER HARNESS 99.9% $18,817 to run Opaque reasoning state carried over Compaction across long contexts THE NUMBER A NEUTRAL RIG PRODUCES THE NUMBER IN THE ANNOUNCEMENT AND THE BETTER SCORE WAS THE CHEAPER RUN 3.66x faster by elapsed time  ·  49% fewer tokens  ·  $7,281 less Scaffolding was not overhead here. It was the capability.
For scale, the same benchmark six months ago: GPT-5.6 Sol scored 7.8%, Claude Opus 5 scored 30.2%. The jump is real under either harness. The headline number is the one that is not reproducible outside OpenAI's own rig.
Data: ARC Prize, September 3, 2026. Chart: nativefirst.ai

What the two harnesses actually differ on

ARC Prize runs a standard harness on purpose: a provider-neutral interface so numbers can be compared across labs. Under it, a model carries forward only the notes it explicitly chooses to keep. Everything it is thinking is visible and portable.

OpenAI's Provider Adapter does something different. It preserves opaque reasoning state between requests and uses compaction to manage long conversations, so the model reuses prior work rather than reconstructing it. That is a legitimate product capability, shipped in the Responses API, and it is genuinely useful. It is also not what the other models in the comparison chart were given.

ARC Prize was careful about the conclusion, and the caution belongs next to the number.

Greg Kamradt
Greg
Kamradt

"Saturating the benchmark would not represent proof of achieving AGI."

ARC Prize · September 3, 2026

The index that did not move

Here is the other number worth holding. On Artificial Analysis's Intelligence Index, Astra scores 61. GPT-5.6 Sol, the model it replaces, also scores 61. Claude Fable 5.1 scores 66.

So an independent aggregate shows no gain at all, on the same day OpenAI published a sweep of state-of-the-art claims. That looks like a contradiction and it is not one. Read the two together and a coherent picture appears:

Astra is a large jump in doing and roughly flat in thinking. The benchmarks it wins measure long-horizon agentic execution: computer use, terminal work, persistence across hours. The index it does not move measures general reasoning quality. Both can be true, and the consequence for you is that this is not a blanket upgrade. It is a task-shaped one.

What people who actually used it said

The early-access reads were unusually consistent: enormous at operating software, less special at judgment.

Dan Shipper
Dan
Shipper

"It's a big upgrade from 5.6-Sol, with some frustrating habits that keep it from matching Fable at the top end."

Dan Shipper, Every · September 3, 2026

Shipper calls it the best writing model he has used, with very little slop and easy steering, and says computer use is "wild, it can go for hours at a time using complicated apps to get work done." Against it: it overcomplicates, and Anthropic's Fable still has better instincts for building a product. The tell is buried in the piece itself. Astra wrote the vibe check's first draft from a single prompt, and Shipper initially mistook it for a colleague's work.

Aaron Levie
Aaron
Levie

"It is now the best model we've ever tested on our expanded and hardest test set."

Aaron Levie, Box, on their enterprise complex-work eval · September 3, 2026

Claire Vo produced the most concrete build log: one-shot coding wins that Fable and 5.6 could not match, an AIM-style Mac desktop app in a single attempt, a hardware CLI she had chased since GPT-5.5, and a ChatPRD feature she had "thrown every model at for six months" finally landing at 90%.

Ethan Mollick pointed it at tens of thousands of his own emails, writings and calendar entries and had it assemble a personal knowledge base of research, contacts, ideas, relationships and tasks. That is the clearest published example of the long-horizon mode, and it is the one most likely to matter inside a company.

The counter-take of the day

Against Brockman's AGI framing, the sharpest reply came from someone describing his own setup rather than the model.

signull
signüll

"All of my personal assistants are now reading each other's updates, and summarizing what the others learned. This feels way more like agi than astra."

@signulll · September 3, 2026

That is the same lesson the ARC-AGI asterisk teaches, arrived at from the opposite direction. In both cases the thing producing the capability is the system around the model, not the weights. A 37-point swing from a harness change and a set of agents summarising each other are two versions of one finding.

What it costs, and what it does not tell you

$10 per million input tokens, $50 per million output. That is 2.5x GPT-5.6 Sol, and it is exactly level with Claude Fable. Note the second comparison, because it is the one that matters for routing: at identical prices, Fable currently sits five points higher on the independent reasoning index, and Astra is far ahead on computer use.

Two billing details that are not in the headline rate. Search and computer-use tool calls bill per call on top of tokens. And rollout started through a limited Daybreak Access programme before ChatGPT tiers, the API and AWS, which Daring Fireball called a soft release and Forbes called a curious false start.

On safety, both frontier labs have now shipped a model with a cyber gate attached. Astra supports defensive work, secure code review and patching, and refuses advanced offensive work such as writing proof-of-concept exploits.

The one thing to take from day one

Not "AGI is here" and not "the benchmarks are fake." Both readings skip the finding.

The most valuable number OpenAI published this week is the gap between 62.7% and 99.9%, because that gap was produced entirely by scaffolding: reasoning state carried between calls, and compaction across a long context. It made the model 37 points better, 3.66x faster and 27% cheaper on the same task set.

If a harness change is worth 37 points to a frontier lab on a benchmark, it is worth more than that to you on your own workflows, where nobody has tuned anything yet. The scaffolding is the product, and it is the part you own.

Buy the harness, not the headline.

Which of your workflows is a harness problem?

Book a free Diagnostic: 30 to 45 minutes, no deck, no pitch. We find the workflow where the model is already good enough and the scaffolding is what is missing, and scope what it takes to close it.

Book the Diagnostic →
Sources
1OpenAI, "GPT-6 Astra: A new generation of intelligence", September 3, 2026: claimed state of the art on FrontierMath Tier 4, ARC-AGI 3, TerminalBench-4.0, Terminal-Bench Science 0.1 and HealthBench Pro; ExploitBench 100% (GPT-5.6 Sol 78.5%); OSWorld 2.0 offline 72.6% (Sol 65.7%, Claude Opus 5 70.2%). Rollout via limited Daybreak Access, then ChatGPT Plus/Pro/Business/Enterprise, the API and AWS. Defensive cyber work supported; advanced offensive work such as proof-of-concept exploits refused. openai.com
2ARC Prize, "OpenAI's GPT-6 Astra on ARC-AGI-3", September 3, 2026: 62.7% on the Semi-Private set at maximum reasoning effort under the Standard harness at a total cost of $26,098, versus 99.9% at high reasoning effort under the Provider Adapter harness at $18,817, which was ~3.66x faster by elapsed time and used 49% fewer tokens. The Provider Adapter "preserves opaque reasoning state between requests and uses compaction for longer conversations, allowing the model to reuse prior work"; the Standard harness carries forward only notes the model chooses to keep. Prior scores: Claude Opus 5 30.2%, GPT-5.6 Sol 7.8%. ARC Prize's caveat: "saturating the benchmark would not represent proof of achieving AGI." arcprize.org
3Simon Willison, "GPT-6 Astra", September 3, 2026, for pricing and the benchmark caveat: $10/million input and $50/million output, matching Claude Fable; "clearly OpenAI's Fable competitor"; on Artificial Analysis's Intelligence Index Astra scores 61, level with GPT-5.6 Sol and behind Claude Fable 5.1 at 66. simonwillison.net
4Sam Altman, September 3, 2026: "We believe it is the best model in the world for computer use, professional work, science, coding, cybersecurity, and more." Greg Brockman's "not unreasonable to feel that we are now in the AGI era" and Tobi Lutke's one-word "Singularity" reaction date from the same day. x.com
5Dan Shipper, "Vibe Check: GPT-6 Astra Is a Big Upgrade With Some Bad Habits", Every, September 3, 2026 (partly paywalled): "a big upgrade from 5.6-Sol, with some frustrating habits that keep it from matching Fable at the top end." Astra wrote the piece's first draft from a single prompt. every.to
6Aaron Levie, September 3, 2026, on Box's enterprise complex-work eval: "It is now the best model we've ever tested on our expanded and hardest test set." x.com
7Claire Vo, "GPT-6 Astra is a banger, here's everything I've built", Lenny's Newsletter, September 3, 2026: one-shot wins including a ChatPRD product-intelligence feature she had failed at for six months, an AIM-style Mac desktop app, a Divoom MiniToo hardware CLI chased since GPT-5.5, and Blender 3D assets. Her rule: browser use "for QA, not building." Also argues "UI is genuinely back" and what that means for SaaS and MCPs. lennysnewsletter.com
8Ethan Mollick, September 3, 2026: assigned GPT-6 tens of thousands of emails, writings and calendar appointments to assemble a personal knowledge base of research, contacts, ideas, relationships and tasks. x.com
9@signulll, September 3, 2026: "all of my personal assistants are now reading each other's updates, & summarizing what the others learned. this feels way more like agi than astra." x.com
10Daring Fireball, "OpenAI Soft-Releases GPT-6 Astra", September 3, 2026, on the staggered rollout. daringfireball.net
John Tan
John Tan

Founder and CEO of nativefirst.ai. Embeds with scaling founders and CEOs to ship Level-3 agents and AI workflows in production.