Since November 2025 this Playbook has tracked three AI benchmarks every month: GDPval-AA for real knowledge work, SWE-bench for coding, and the Artificial Analysis Intelligence Index for general capability. Nineteen posts, eleven months, one leaderboard after another collapsing under its own weight.
This is the index. Every update in one place, plus the three explainers that tell you what the numbers mean before you quote one at a board.
Where to start
Start here: what the numbers mean
Three explainers, written once and still current. Read the one that matches the decision in front of you.
- What Is GDPval? The benchmark that scores models on real deliverables from 44 occupations, graded by working professionals. The one to quote if you are buying AI for knowledge work rather than code.
- What Is SWE-bench? Real GitHub issues, real repositories, a passing test suite or nothing. The coding benchmark, and the reason its Verified split had to be retired.
- What Is the Artificial Analysis Intelligence Index? The composite that most launch posts cite without naming. What goes into it, and what happened when it was rebuilt.
The combined update: all benchmarks, plus price
From July 2026 the two monthly trackers merged into one. The reason was not editorial tidiness. It was that price had become part of the leaderboard, and a coding score without a cost next to it had stopped being a buying signal.
- Benchmark Update (August 2026): Grok 4.6 Joins the Frontier. Opus 5 at 63, Fable 5 at 62, Grok 4.6 and Sol at 61 on the AA index, with prices spanning 8x across four models separated by two points.
- Benchmark Update (July 2026): Opus 5 Matches Fable at Half. The busiest benchmark month of the year, and the arrival of the priced leaderboard. Opus 5 hit 66.7 against Fable's 66.5 on CursorBench at half the cost.
- AI Benchmarks Update (January 2026): The Index Overhaul. Artificial Analysis rebuilt the Intelligence Index: ten work-shaped evals in, saturated exams out, top scores down 20 points overnight.
GDPval: real work, graded by professionals
Seven updates. Watch the leader change hands five times and the Elo scale get reset underneath everyone once.
- August 2026: Opus 5 Leads at 1862 Elo. GDPval-AA moved to v2 and reset every June number. Grok 4.6 sits at 1753, and xAI's number-one claim does not hold.
- June 2026: The Benchmark That Actually Matters. Claude Fable 5 leads at 1932 Elo. What expert parity on real deliverables means, and the catch in the fine print.
- May 2026: The Leaderboard Reshuffles. Opus 4.8 takes the lead at 1890 Elo, Grok 4.3 jumps 321 points, and Gemini 3.5 Flash beats Google's own Pro tier on real work.
- April 2026: GPT-5.5 Sets the Bar. Launches at 84.9% expert parity, and economists start writing about AI eating analyst work.
- March 2026: GPT-5.4 Crowds the Top. 1674 Elo, 41 points above Sonnet 4.6. Three labs within 70 points, and what that compression means for buyers.
- February 2026: Anthropic Takes Both Top Slots. Sonnet 4.6 at 1633 Elo and Opus 4.6 at 1606, with Gemini 3.1 Pro last among the majors.
- December 2025: The Leaderboard Arrives. GPT-5.2 hits 70.9% win or tie against professionals, and Artificial Analysis launches GDPval-AA as an independent Elo leaderboard.
SWE-bench: code that has to compile
Six updates, and the slowest-motion benchmark death in the set. Verified was deprecated in February, called a zombie in May, and killed in June.
- June 2026: Fable 5 Tops a Dying Leaderboard. 80.3% on SWE-bench Pro, OpenAI kills Verified for good, and FrontierCode resets the leaderboard underneath.
- May 2026: Opus 4.8 Takes the Lead. 69.2% on Pro, and open-weights models close to within 6 points at 8x lower cost.
- April 2026: The Month the Benchmark Broke. Berkeley researchers break eight agent benchmarks, and Mythos Preview exposes the Verified-versus-Pro gap.
- March 2026: GPT-5.4 Takes Pro. 59.1% on SEAL's standardized Pro split, with Opus 4.6 holding the commercial subset.
- February 2026: The Month Verified Died. OpenAI deprecates Verified over contamination, and Pro takes over as the number that counts.
- November 2025: Opus 4.5 Breaks 80. The first model over 80% on Verified, after four frontier releases in twelve days. Pro sat at 45.9%.
What eleven months of scores actually show
The benchmark broke four times. The models did not stop improving. The Intelligence Index was rebuilt in January and top scores fell 20 points overnight. Verified was deprecated in February over contamination. Berkeley broke eight agent benchmarks in April. GDPval-AA moved to v2 in August and reset every June number. Through all of it, SWE-bench Pro went from 45.9% in November 2025 to 80.3% in June 2026. The instrument kept failing and the thing being measured kept moving.
No lab held the lead for two consecutive updates. GPT-5.2, then Sonnet 4.6, then GPT-5.4, then GPT-5.5, then Opus 4.8, then Fable 5, then Opus 5. Seven leaders in nine months. If you picked a model because it was top of a leaderboard, you have been wrong roughly every six weeks, and the cost of being wrong has been near zero. That is the argument for routing rather than standardising.
The top of the table compressed until price became the tiebreak. In March, three labs sat within 70 Elo points. By August, four models sat within two points on the AA index with prices spanning 8x. A two-point capability gap and an 8x price gap is not a capability decision. It is a procurement decision, and that is why the July update started publishing cost next to score.
Open weights stopped being a separate conversation. By May 2026 open-weights models were within 6 points of the frontier on SWE-bench Pro at 8x lower cost. If you have not priced that option since then, you are working from a stale number. Start with What Are Open Weights?
The caveat that belongs on every number here
Every figure in this index is a published score, and published scores are marketing before they are measurement. Labs choose which benchmark to lead with, contamination is real and was the stated reason Verified died, and a reset means last month's number is not comparable to this month's even when both are on the page.
Read The Benchmarks Are Getting Gamed before you put any of these numbers in a deck. The short version: use benchmarks to eliminate models, not to pick one. The picking is done on your own work, with your own evals, which is the only test that cannot be trained against.
Updated monthly. The September 2026 edition is next.