Everything in this series is real, and you have now watched most of it. Nothing below retracts that.

But the builds share a set of properties, and once you see the set you cannot unsee it. Each was made by an expert practitioner. Each had early access. Each was a task the builder chose. Each is one attempt, shown once, by the person who made it, with nobody downstream depending on the result.

That is not a criticism. It is the definition of a demo. The problem is what happens when a company watches twenty of them and concludes it has seen a product.

The gap, itemised

Fig. 1
What sits between a demo and a system
Everything the clip does not have to do WHAT A DEMO NEEDS one expert operator one attempt, the one you show a task picked because it works no error path no cost ceiling no second user 6 things WHAT A SYSTEM ALSO NEEDS an eval set that runs on every change a defined path when it fails a cost ceiling per run credential scope and blast radius someone on call a user who is not the builder 6 more, and these are the hard ones LAUNCH WEEK PUBLISHED THE LEFT COLUMN. THE RIGHT COLUMN IS THE JOB.
None of this makes the builds fake. It makes them the first six of twelve.

Read the right-hand column, because that is the work. An eval set that runs on every change is the difference between "it worked in September" and "it still works in November." A defined failure path is the difference between an agent that stops and an agent that improvises with your production data. A cost ceiling is the difference between a $40 run and a $4,000 run you find on the invoice.

None of it appears in a clip, because none of it is visible in one. It is also, in my experience, four fifths of the actual build.

The number the wave did not post

Artificial Analysis Intelligence Index

61
GPT-5.6 Sol
61
GPT-6 Astra
66
Claude Fable 5.1
2.5x
The price increase

An independent aggregate showed no gain at all against the model Astra replaced, at two and a half times the price, from a launch claiming state of the art on five benchmarks.

Both readings hold. Astra jumped hard on agentic execution and long-horizon persistence, which is what every build in this series exercises. It did not move on general reasoning. The demos are demos of doing, and the doing is genuinely better.

Even the builders said so

The 3D teardown in Part 1 found character modelling "requires an exhausting amount of manual micro-management." The design comparison in Part 2 declined to call Astra better overall and caught a wrong domain rendered on screen. Yunfan Ye flagged "still some wrong details" in his own post.

Dan Shipper
Dan
Shipper

"A big upgrade from 5.6-Sol, with some frustrating habits that keep it from matching Fable at the top end. It overcomplicates, and Fable still has better instincts for building a product."

Every, Vibe Check · September 3, 2026

And the sharpest counter to the AGI framing came the same day it was made:

signull
signull

"All of my personal assistants are now reading each other's updates, and summarizing what the others learned. This feels way more like agi than astra."

That is the right instinct and the thesis of this whole site. The interesting thing was not the model. It was several agents arranged into something that held state and passed work between parts. The system did the work.

Then the rollout broke

Sam Altman
Sam
Altman

"First, sorry for the messy rollout. Second, when we screw up, we try to make it right. Third, we should be able to begin broad rollout to API customers and chatgpt subscribers in the near future."

Worth holding against the gallery: for most of launch week, most people could not run any of it.

So what do you do with twenty demos

Treat each as a hypothesis. "Astra can map an undocumented system" is testable against your own undocumented system in an afternoon. Do that instead of believing it.

Route rather than migrate. Send computer use, 3D, interface work and long-horizon agentic tasks to Astra. Leave general reasoning where it is, where it is currently better and cheaper.

Count the right column. Before anything here goes near production, write down the eval set, the failure path, the cost ceiling, the credential scope and the name of the person on call. If you cannot fill in all five, you have a demo.

The wave was real. It was also six of twelve.

Turn one of these into a system.

Book a free Diagnostic: 30 to 45 minutes, no deck, no pitch. We pick one pattern from the gallery and scope what production would actually take.

Book the Diagnostic →

Part 4 of Astra, Week One.

Sources
1Simon Willison, "GPT-6 Astra", September 3, 2026: on Artificial Analysis's Intelligence Index Astra scores 61, level with GPT-5.6 Sol and behind Claude Fable 5.1 at 66; pricing $10 per million input and $50 per million output. simonwillison.net
2Every, "Vibe Check: GPT-6 Astra Is a Big Upgrade With Some Bad Habits", September 3, 2026, byline Katie Parrott. every.to
3Sam Altman (@sama), September 4, 2026, on the rollout; Tibo Sottiaux the same day offering one banked reset per day without access. Compensation denominated in usage, not money.
4Stated weaknesses quoted from yage.ai and mindstudio.ai; Yunfan Ye's caveat is in his own embedded post in Part 1.
John Tan
John Tan

Founder and CEO of nativefirst.ai. Embeds with scaling founders and CEOs to ship Level-3 agents and AI workflows in production.