This series went quiet after June, and July turned out to be the busiest benchmark month of the year: three frontier launches, one open-weights arrival, and the first leaderboard where the interesting column is price.

The June update ended on a dying leaderboard, with Fable 5 topping SWE-bench Pro at 80.3% while OpenAI killed Verified for contamination. July's story is what replaced single-number rankings: cost per completed task.

The month in three launches

July 9: GPT-5.6 Sol, Terra and Luna. OpenAI's launch numbers put Sol at 53.6 on Agents' Last Exam, 13.1 points clear of Fable 5 adaptive. Vendor-published, never independently replicated at that margin, but the model held up in a month of field use.

July 16: Kimi K3. Moonshot's 2.8T open-weights model arrived one point-cluster below the frontier at a fraction of the price, and reset what the bottom half of every leaderboard costs.

July 24: Claude Opus 5. The quiet headline of the month. On CursorBench 3.2, Cursor measured Opus 5 at 66.7 against Fable 5's 66.5 at default effort, at half the price and with zero-data-retention compatibility Fable lacks.

The leaderboard that matters now

CursorBench 3.2's full table, published July 24, is the first mainstream board to price every row. Read it as score against dollars per completed task:

Fig. 1
Score against cost per task, July 2026
CursorBench 3.2, priced SCORE · $/TASK JULY 24, 2026 Fable 5 Max 70.5% $17.32 Opus 5 Max 70.0% $8.23 Sonnet 5 Max 61.5% $6.45 Luna Max 61.1% $1.97 Kimi K3 Max 60.8% $2.70 Kimi K3 High 59.7% $1.89 Half a point separates the top two at a 2.1x price gap. Ten points separate top from bottom at 9.2x.
The score column has converged. The price column has not. Routing decisions now live in the right-hand column.
Source: CursorBench 3.2 public table, July 24, 2026

Two readings matter. At the top, Opus 5 within half a point of Fable at less than half the cost per task is why the frontier premium argument got harder to make in July. In the middle band, four models from three labs sit within two points of each other across a nine-times price range, which is the entire case for routing by workload.

Also in July

The first open-model IMO gold. Nvidia's open Nemotron 3 Ultra scored gold-medal level at the International Mathematical Olympiad, 30 of 42 as graded by the IMO team, with Ethan Mollick noting Kimi K3 and GLM-5.2 would likely also qualify.

The Harbor Town split. The one-shot gallery scoring code and design separately became the most useful single chart of the month: Opus 5 at 20/20 on both, Kimi K3 at 19/20 code but 14/20 design. Open models reached code parity in July. They did not reach taste parity, and single-number leaderboards hide the difference.

A benchmark caveat that aged well. Every headline number this month is either vendor-published or single-source. The gaming problem from June did not go away; it moved into the launch decks. Treat any un-replicated margin as marketing until a second measurement lands.

July's summary in one line: the score gap at the frontier closed to noise, and the price column became the leaderboard.

Route by the right column.

Book a free Diagnostic: 30 to 45 minutes, no deck, no pitch. We work out what your workloads cost per completed task, and which rows of this table they should be running on.

Book the Diagnostic →
Sources
1Cursor (@cursor_ai), July 24, 2026: Claude Opus 5 measured at 66.7 vs Fable 5's 66.5 on CursorBench at default effort, at half the price and ZDR-compatible. x.com
2CursorBench 3.2 public table, July 24, 2026, via @nateschmidt: Fable 5 Max 70.5% at $17.32/task; Opus 5 Max 70.0% at $8.23; Sonnet 5 Max 61.5% at $6.45; GPT-5.6 Luna Max 61.1% at $1.97; Kimi K3 Max 60.8% at $2.70; Kimi K3 High 59.7% at $1.89. x.com
3OpenAI launch figures, July 9, 2026: GPT-5.6 Sol at 53.6 on Agents' Last Exam, +13.1 over Claude Fable 5 adaptive. Vendor-published; not independently replicated at that margin.
4Ethan Mollick (@emollick), July 22, 2026: Nvidia's open Nemotron 3 Ultra at gold-medal level on the IMO, 30/42 graded by the IMO team; Kimi K3 and GLM-5.2 would likely also qualify. x.com
5Harbor Town one-shot gallery, July 2026 scores: Opus 5 20/20 code and 20/20 design; Kimi K3 19/20 and 14/20. ai-harbor-town-gallery.netlify.app
6Moonshot Kimi K3 release, July 16, 2026, covered in full in the Open Weights series. nativefirst.ai
John Tan
John Tan

Founder and CEO of nativefirst.ai. Embeds with scaling founders and CEOs to ship Level-3 agents and AI workflows in production.