This series went quiet after June, and July turned out to be the busiest benchmark month of the year: three frontier launches, one open-weights arrival, and the first leaderboard where the interesting column is price.
The June update ended on a dying leaderboard, with Fable 5 topping SWE-bench Pro at 80.3% while OpenAI killed Verified for contamination. July's story is what replaced single-number rankings: cost per completed task.
The month in three launches
July 9: GPT-5.6 Sol, Terra and Luna. OpenAI's launch numbers put Sol at 53.6 on Agents' Last Exam, 13.1 points clear of Fable 5 adaptive. Vendor-published, never independently replicated at that margin, but the model held up in a month of field use.
July 16: Kimi K3. Moonshot's 2.8T open-weights model arrived one point-cluster below the frontier at a fraction of the price, and reset what the bottom half of every leaderboard costs.
July 24: Claude Opus 5. The quiet headline of the month. On CursorBench 3.2, Cursor measured Opus 5 at 66.7 against Fable 5's 66.5 at default effort, at half the price and with zero-data-retention compatibility Fable lacks.
The leaderboard that matters now
CursorBench 3.2's full table, published July 24, is the first mainstream board to price every row. Read it as score against dollars per completed task:
Two readings matter. At the top, Opus 5 within half a point of Fable at less than half the cost per task is why the frontier premium argument got harder to make in July. In the middle band, four models from three labs sit within two points of each other across a nine-times price range, which is the entire case for routing by workload.
Also in July
The first open-model IMO gold. Nvidia's open Nemotron 3 Ultra scored gold-medal level at the International Mathematical Olympiad, 30 of 42 as graded by the IMO team, with Ethan Mollick noting Kimi K3 and GLM-5.2 would likely also qualify.
The Harbor Town split. The one-shot gallery scoring code and design separately became the most useful single chart of the month: Opus 5 at 20/20 on both, Kimi K3 at 19/20 code but 14/20 design. Open models reached code parity in July. They did not reach taste parity, and single-number leaderboards hide the difference.
A benchmark caveat that aged well. Every headline number this month is either vendor-published or single-source. The gaming problem from June did not go away; it moved into the launch decks. Treat any un-replicated margin as marketing until a second measurement lands.
July's summary in one line: the score gap at the frontier closed to noise, and the price column became the leaderboard.
Route by the right column.
Book a free Diagnostic: 30 to 45 minutes, no deck, no pitch. We work out what your workloads cost per completed task, and which rows of this table they should be running on.
Book the Diagnostic →