Since November 2025 this Playbook has tracked three AI benchmarks every month: GDPval-AA for real knowledge work, SWE-bench for coding, and the Artificial Analysis Intelligence Index for general capability. Nineteen posts, eleven months, one leaderboard after another collapsing under its own weight.

This is the index. Every update in one place, plus the three explainers that tell you what the numbers mean before you quote one at a board.

Figure 1. The record
Nineteen posts, three benchmarks, four resets
GDPval-AA SWE-bench Combined + price next NOVDECJAN FEBMARAPR MAYJUNJUL AUGSEP RESETS Two trackers, running in parallel Merged, and priced INDEX REBUILT  /  VERIFIED DEPRECATED  /  8 BENCHMARKS BROKEN  /  GDPVAL-AA V2
Four of the eleven months contained a reset that invalidated earlier numbers. That is the single most useful fact on this page.

Where to start

New to AI benchmarks123
You just need the current numbers45
You are buying a coding model26
You are buying for knowledge work15
Someone quoted a benchmark at you73
Start from zero31247

Start here: what the numbers mean

Three explainers, written once and still current. Read the one that matches the decision in front of you.

The combined update: all benchmarks, plus price

From July 2026 the two monthly trackers merged into one. The reason was not editorial tidiness. It was that price had become part of the leaderboard, and a coding score without a cost next to it had stopped being a buying signal.

GDPval: real work, graded by professionals

Seven updates. Watch the leader change hands five times and the Elo scale get reset underneath everyone once.

SWE-bench: code that has to compile

Six updates, and the slowest-motion benchmark death in the set. Verified was deprecated in February, called a zombie in May, and killed in June.

What eleven months of scores actually show

The benchmark broke four times. The models did not stop improving. The Intelligence Index was rebuilt in January and top scores fell 20 points overnight. Verified was deprecated in February over contamination. Berkeley broke eight agent benchmarks in April. GDPval-AA moved to v2 in August and reset every June number. Through all of it, SWE-bench Pro went from 45.9% in November 2025 to 80.3% in June 2026. The instrument kept failing and the thing being measured kept moving.

No lab held the lead for two consecutive updates. GPT-5.2, then Sonnet 4.6, then GPT-5.4, then GPT-5.5, then Opus 4.8, then Fable 5, then Opus 5. Seven leaders in nine months. If you picked a model because it was top of a leaderboard, you have been wrong roughly every six weeks, and the cost of being wrong has been near zero. That is the argument for routing rather than standardising.

The top of the table compressed until price became the tiebreak. In March, three labs sat within 70 Elo points. By August, four models sat within two points on the AA index with prices spanning 8x. A two-point capability gap and an 8x price gap is not a capability decision. It is a procurement decision, and that is why the July update started publishing cost next to score.

Open weights stopped being a separate conversation. By May 2026 open-weights models were within 6 points of the frontier on SWE-bench Pro at 8x lower cost. If you have not priced that option since then, you are working from a stale number. Start with What Are Open Weights?

The caveat that belongs on every number here

Every figure in this index is a published score, and published scores are marketing before they are measurement. Labs choose which benchmark to lead with, contamination is real and was the stated reason Verified died, and a reset means last month's number is not comparable to this month's even when both are on the page.

Read The Benchmarks Are Getting Gamed before you put any of these numbers in a deck. The short version: use benchmarks to eliminate models, not to pick one. The picking is done on your own work, with your own evals, which is the only test that cannot be trained against.

Updated monthly. The September 2026 edition is next.

Sources
1Artificial Analysis, GDPval-AA leaderboard and Intelligence Index, November 2025 through August 2026. GDPval-AA moved to v2 in August 2026, resetting prior Elo figures.
2OpenAI, GDPval, covering real deliverables across 44 occupations graded by working professionals.
3SWE-bench (Princeton) and SWE-bench Pro, including OpenAI's deprecation of the Verified split in February 2026 over contamination and its retirement in June 2026.
4Scale AI SEAL leaderboard, standardized SWE-bench Pro results, March 2026.
5UC Berkeley agent-benchmark analysis covering eight agent benchmarks, April 2026.
6All figures are published scores from the labs and independent leaderboards named in each linked update, recorded on the date of that update. A benchmark reset means figures either side of it are not directly comparable.
John Tan
John Tan

Founder and CEO of nativefirst.ai. Embeds with scaling founders and CEOs to ship Level-3 agents and AI workflows in production.