A week ago the frontier was a two-horse race and everybody knew it. Anthropic and OpenAI, trading the top of the board between them, with xAI somewhere behind and fading.
Austen Allred, who had said exactly that, on August 12:
Allred
It's insane how xAI was basically not a contender and now they're right back in the race. I honestly can't believe it. I thought they were toast and it was a two-horse race.
Here is what changed his mind.
The board, as measured
Artificial Analysis put Grok 4.6 at 61 on its Intelligence Index. For context on the same index: Claude Opus 5 sits at 63, Claude Fable 5 at 62, GPT-5.6 Sol at 61.
So Grok 4.6 is one point behind Fable 5, two behind Opus 5, and level with Sol. That alone would be a story, because Grok 4.5 was five points lower a month ago.
The price is the actual story. Grok 4.6 bills at $2 per million input tokens and $6 per million output. Fable 5 bills at $10 and $50.
Five weeks
The speed of it is easy to miss. In early July, Grok 4.5 arrived and was described as Opus-class: fast, token-efficient, competitive but not leading. On August 11, xAI shipped Grok Bot, agents with their own cloud computer that sign in to your tools and finish jobs. On August 12, Grok 4.6 landed at the frontier.
Underneath that sits an unusual amount of infrastructure. SpaceX acquired Anysphere, Cursor's parent, for roughly $60B in June, which is how Grok Bot arrived attached to Cursor's distribution. And Grok 4.6 gained five index points in a month, which is a rate of improvement nobody was pricing in.
Elon has already said Grok 4.7 will exceed all current models, which is what he says before every release and is therefore worth precisely nothing as evidence. The measured number is the one to hold onto.
This contradicts something I published two days ago
On Tuesday this playbook argued that the floor collapsed and the frontier did not. The chart in that post showed cheap models falling to near-zero while flagship pricing held around $5 to $10 per million input, roughly where GPT-4 launched.
Grok 4.6 is a model one point off the top of the board at $2 in and $6 out. That is not the floor falling. That is the frontier getting cheap.
The argument in that post still holds in its conclusion, which is that routing matters more than rates. But its central framing needs an amendment, and it is better to say so here than to let it quietly age. Two days is a short shelf life for a thesis. That is the actual condition of this market and pretending otherwise helps nobody.
Aaron Levie, after DeepSeek and Grok both shipped inside the same week: this is truly Jevons paradox for AI.
Updated the Same Evening: The Field Reports
This post went up on the morning of August 13 arguing from the index. By evening the field reports had landed, and they are better evidence than any benchmark, so here they are.
The one that matters most is DHH. A few days earlier he had shown Claude Fable 5 one-shotting a Rust rewrite of the TerminalTextEffects Python library in 11M tokens: startup time from 87ms to 2ms, rendering 9.6x faster, zero dependencies, a 3MB single executable. Then he handed Fable's plan to Grok 4.6:
I had @spacexai Grok 4.6 follow Fable's plan, and with just a couple of nudges, it was able to repeat this feat in 1h 24m using 8.6M tokens at a cost of ~$55 at per-token pricing. That's about 1/10 the cost of the Fable implementation for the same work!
That is the Fig. 1 gap, reproduced on a real codebase by someone with no stake in the answer. Same work, a tenth of the bill. Note the asterisk that also matters: Grok followed Fable's plan. The expensive model did the thinking, the cheap one did the labor. That split is a routing strategy, and it is exactly where this is heading.
The fuller stress test came from Kun Chen, who ran Grok 4.6 through a full day of real work: as the orchestrator agent he talks to directly, managing all his other agents, and as a worker agent on implementation tasks. His review opens with the good, starting with speed. A day of orchestrator duty is a harder exam than any eval suite.
And the adoption signal: Claire Vo observed that more and more people she knows are skipping the Anthropic/OpenAI debate entirely and becoming grok bois instead. When the default question stops being which of the two, the race has a third horse by definition.
Now the honesty paragraph. On DeepSWE v1.1, Grok 4.6 scores 65.9% against 73% for GPT-5.6 Sol Max. On the wider ten-row comparison table circulating in the model coverage, Claude Fable 5 still wins the most rows. And no independent third party has replicated xAI's own numbers yet. So hold both things at once: the arrival is real, the crown claims are not.
What a third lab actually changes
Not capability. One index point is inside the noise of any real workload, and you should distrust anyone ranking their own model, including all three of these companies.
What changes is procurement.
Single-vendor architecture is now a choice, not a constraint. For two years there was a defensible argument that you built on one lab because only one or two could do the work. Three labs at the frontier, with the cheapest at a fifth of the price, ends that argument. If your system cannot swap models, that is now a decision you made rather than a limitation you inherited.
The switching test is worth running. One indie developer noted this week that Claude had become preachy enough that he was ready to move to Grok for coding, and that his sites already run on xAI's backend. Whether or not he is right about the model, the fact that switching is casually discussable is new.
Price pressure now runs upward. When the cheap tier got cheaper, frontier labs could ignore it. When a near-frontier model prices at a fifth, they cannot. Expect the response within a quarter.
The lesson from the Fable 5 episode was to rent the model and own the loop. Nothing about a third contender changes that. It just makes the renting cheaper and the owning more valuable, because the thing you own is the only part that does not get repriced every five weeks.
Rent the model. Own the loop.
Can your stack swap models?
Book a free Diagnostic: 30 to 45 minutes, no deck, no pitch. We look at what you have built on top of a single model, and what it would take to move it.
Book the Diagnostic →