Everything in this series is real, and you have now watched most of it. Nothing below retracts that.
But the builds share a set of properties, and once you see the set you cannot unsee it. Each was made by an expert practitioner. Each had early access. Each was a task the builder chose. Each is one attempt, shown once, by the person who made it, with nobody downstream depending on the result.
That is not a criticism. It is the definition of a demo. The problem is what happens when a company watches twenty of them and concludes it has seen a product.
The gap, itemised
Read the right-hand column, because that is the work. An eval set that runs on every change is the difference between "it worked in September" and "it still works in November." A defined failure path is the difference between an agent that stops and an agent that improvises with your production data. A cost ceiling is the difference between a $40 run and a $4,000 run you find on the invoice.
None of it appears in a clip, because none of it is visible in one. It is also, in my experience, four fifths of the actual build.
The number the wave did not post
Artificial Analysis Intelligence Index
An independent aggregate showed no gain at all against the model Astra replaced, at two and a half times the price, from a launch claiming state of the art on five benchmarks.
Both readings hold. Astra jumped hard on agentic execution and long-horizon persistence, which is what every build in this series exercises. It did not move on general reasoning. The demos are demos of doing, and the doing is genuinely better.
Even the builders said so
The 3D teardown in Part 1 found character modelling "requires an exhausting amount of manual micro-management." The design comparison in Part 2 declined to call Astra better overall and caught a wrong domain rendered on screen. Yunfan Ye flagged "still some wrong details" in his own post.
Shipper
"A big upgrade from 5.6-Sol, with some frustrating habits that keep it from matching Fable at the top end. It overcomplicates, and Fable still has better instincts for building a product."
And the sharpest counter to the AGI framing came the same day it was made:
"All of my personal assistants are now reading each other's updates, and summarizing what the others learned. This feels way more like agi than astra."
That is the right instinct and the thesis of this whole site. The interesting thing was not the model. It was several agents arranged into something that held state and passed work between parts. The system did the work.
Then the rollout broke
Altman
"First, sorry for the messy rollout. Second, when we screw up, we try to make it right. Third, we should be able to begin broad rollout to API customers and chatgpt subscribers in the near future."
Worth holding against the gallery: for most of launch week, most people could not run any of it.
So what do you do with twenty demos
Treat each as a hypothesis. "Astra can map an undocumented system" is testable against your own undocumented system in an afternoon. Do that instead of believing it.
Route rather than migrate. Send computer use, 3D, interface work and long-horizon agentic tasks to Astra. Leave general reasoning where it is, where it is currently better and cheaper.
Count the right column. Before anything here goes near production, write down the eval set, the failure path, the cost ceiling, the credential scope and the name of the person on call. If you cannot fill in all five, you have a demo.
The wave was real. It was also six of twelve.
Turn one of these into a system.
Book a free Diagnostic: 30 to 45 minutes, no deck, no pitch. We pick one pattern from the gallery and scope what production would actually take.
Book the Diagnostic →Part 4 of Astra, Week One.