GPT-6 Astra shipped on September 3. What made launch week different was not the benchmark table. It was that the people with early access stopped posting scores and started posting things they had made, in professional software, from prompts that were often a single sentence.

This page is the gallery. Every clip below is the original post, playable here. Then, at the end, the number nobody put in a thread.

20
Builds
collected
3,295
Editable objects
from one sketch
15 → 1
Iterations Fable needed
vs Astra's first draft
61
Intelligence index
GPT-5.6 Sol also 61

Blender and 3D

The biggest category, and the one where the artifacts are hardest to argue with. Astra operates Blender rather than telling you how to.

3,295 editable objects from one drawing@tomkrcha
Modelled, rigged with 50 bones, then playable in Unreal@mreflow
The entire prompt was "Make this in Blender"@skirano
Listing photos to a 3D house and a promo video@realYunfanYe
Blender house carried into Unreal Engine 5@goofyninjaaa
The same villa brief, Astra against Claude Fable 5.1@karankendre

Part 1 takes these apart, including what it is still bad at.

Interfaces

The quietest category and possibly the most commercially relevant.

OpenAI shipped front-end design as a stated capability@OpenAIDevs
A 3D iPod, built in Blender, browsing Codex threads@skirano
Real tickets finished, verified, and closed in Linear@mehulmpt
Full colour grades in Final Cut Pro, by clicking@davis7

Part 2 covers the 15-iteration comparison and why flatness was always the tell.

The specialist demos

The three worth showing a sceptical executive, because in each case the benchmark was a person with decades of domain expertise.

A five-minute T cell lecture from one sentence@DeryaTR_
Museum mapped, cast blocked, shot list planned@higgsfield_ai

The third is not on X: Claire Vo reverse-engineered a closed Bluetooth speaker with no API, covered in Part 3.

Games, worlds and simulation

Unreal world of Astra-powered agents that started talking@mattshumer_
An FPS map, built the same morning@rileybrown
A Fall Guys clone@MatthewBerman
A browser kart racer with 8 racers, drifting and items@goofyninjaaa

One correction worth making, because the roundups got it wrong: the kart racer is Pietro Schirano's, built with Sites inside ChatGPT. Goofy is describing it, in Spanish, not claiming it.

Graphics and video

Six Van Gogh paintings merged into one walkable town@petergostev
An open ocean simulator extended with animal behaviour@emollick
A neo-gothic drowned city, as a shader@emollick
And the official number under all of it: 72.6% on OSWorld 2.0@OpenAIDevs

Why the index did not move

Hold all of that against one number and the week resolves.

Fig. 1
Where the jump actually landed
It moved on doing. Not on thinking. DOING  /  AGENTIC EXECUTION ExploitBench78.5% → 100% GPT-5.6 Sol to GPT-6 Astra OSWorld 2.0 offline65.7% → 72.6% Claude Opus 5 sits between, at 70.2% Zapier AutomationBenchNEW RECORD Third party, millions of live workflows THINKING  /  GENERAL REASONING Artificial Analysis Intelligence Index GPT-5.6 Sol61 GPT-6 Astra61 Claude Fable 5.166 Change against the prior flagship 0 points THE BENCHMARKS IT WINS MEASURE DOING. THE INDEX IT DOES NOT MOVE MEASURES THINKING.
A coherent picture, not a contradiction. It just means the model choice is task-shaped, not a blanket upgrade.

Astra jumped hard on agentic execution and long-horizon persistence, which is exactly what every build above exercises. It is flat on general reasoning. The benchmarks it wins measure doing. The index it does not move measures thinking.

The four parts

Which of these maps to a job you actually have?

Book a free Diagnostic: 30 to 45 minutes, no deck, no pitch. We work out which pattern here is worth building in your company, and which is a party trick.

Book the Diagnostic →
Sources
1Every embedded post above is the primary source for its own claim, and each was verified against X before publication. Posts by @tomkrcha, @mreflow, @skirano, @realYunfanYe, @goofyninjaaa, @karankendre, @OpenAIDevs, @mehulmpt, @davis7, @DeryaTR_, @higgsfield_ai, @mattshumer_, @rileybrown, @MatthewBerman, @petergostev and @emollick, September 3 to 4, 2026.
2Simon Willison, "GPT-6 Astra", September 3, 2026, for the independent aggregate: Artificial Analysis Intelligence Index 61, level with GPT-5.6 Sol and behind Claude Fable 5.1 at 66. simonwillison.net
3Claire Vo, How I AI, September 2026, for the Divoom speaker, the AIM-style Mac app and the Figma run. chatprd.ai
4MindStudio for the 15-iteration web-design comparison (mindstudio.ai) and yage.ai for the 3D teardown and its stated weaknesses (yage.ai).
5OpenAI, "GPT-6 Astra: A new generation of intelligence", September 3, 2026. openai.com
John Tan
John Tan

Founder and CEO of nativefirst.ai. Embeds with scaling founders and CEOs to ship Level-3 agents and AI workflows in production.