On September 17, Ramp Labs published a benchmark it built with working accountants. 137 tasks, each with a grading rubric, grounded in everyday accounting workflows. Then it ran the frontier models against them.

The model that finished the most tasks was not from OpenAI and not from Anthropic. It was GLM-5.3, an open-weights model from Z.ai, and it fully solved 21.2% of them with up to three attempts.

Both halves of that sentence matter. An open model beat both US frontier labs at finishing real professional work. And the winner failed four times out of five.

Figure 1. Ramp Accounting Bench
Two metrics, two different winners
FULLY SOLVED Led by GLM-5.3, open weights, three attempts WITH PARTIAL CREDIT Led by Claude Fable 5.1, average score 21.2% 49.7% 29 OF 137 FINISHED HALF MARKS ON MOST OF THEM The metric you pick picks the winner.
A half-correct journal entry is a wrong journal entry. Only the left-hand picture describes work you could file.

The first time open weights led on finishing, not on price

The open-weights argument has spent a year resting on two claims: benchmark parity at a fraction of the cost, and sovereignty. Both are arguments about value. Neither is an argument that an open model is the best tool for a job.

This is different. On the hardest metric in a practitioner-built benchmark, an open model came first outright, ahead of Claude Fable 5.1 and GPT-6 Astra. Not cheaper for almost as good. Better at the thing being measured.

Hold the size of that claim honestly: it is one benchmark, in one profession, published by one company. It is not a general capability ranking and Z.ai has not taken the frontier. But it is the first datapoint in this Playbook where the open option wins on completion rather than on the price of getting close, and that is a different conversation with a CFO than the one we have been having. See What Are Open Weights? and the Kimi K3 gap.

Why "finished" is the only metric that counts here

Ramp published two numbers, and conflating them is the easiest way to misread the whole thing.

Average score with partial credit. Claude Fable 5.1 led at 49.7%. GPT-6 Astra came in at 48.7%, which Ramp notes was achieved "at roughly one third the cost." On this metric the frontier looks halfway there.

Fully solved tasks, up to three attempts each. 21.2%, and the leader changes.

In most domains that distinction is academic. In accounting it is the whole thing. A journal entry that is half right is not half filed. It is wrong, somebody has to find it, and the finding costs more than the entry did. Half marks measure how close a model got. Finished work measures whether anyone can stop watching.

Ramp Labs drew the conclusion itself, in its own launch post: "Reliable agentic accounting remains an open problem." That sentence was written by a company that sells finance automation, which is why it is worth more than any vendor demo you will see this quarter.

Why this benchmark is citable when a general one is not

Three properties, and all three matter for whether you can put a number in front of a board.

It was built by practitioners of the domain, not by a lab deciding what accounting looks like. It is rubric-graded rather than answer-matched, so it tolerates the many-right-answers reality of real work. And it is not in the training corpus, which is the property quietly killing every general benchmark. See The Benchmarks Are Getting Gamed.

IT
Ian
Tracey

"The world needs more vertical benchmarks that aren't exposed in training data. What's notable is even Fable and Astra can only fully solve ~20% of the tasks. There's a lot of room to run."

X  ·  September 18, 2026

That is the same argument GDPval made for knowledge work, narrowed to one regulated profession where the answer is checkable and the cost of being wrong is legible.

One day earlier, a bank said the opposite

On September 16, Mercury launched Mercury Books, which Immad Akhund described as software that "categorizes and reconciles your transactions the moment they happen." Ramp's benchmark landed the following day.

They are not strictly the same task. Categorising transactions against a ledger the bank already owns is narrower and better conditioned than 137 general accounting workflows, and Mercury's structural advantage is real: collapse the bank and the books into one record and reconciliation stops existing rather than being automated.

But owning the ledger removes the drift problem, not the judgment problem, and the judgment problem is what Ramp measured at 21%. A bank that owns every transaction still has to decide what each one means.

The human bar, for scale

The Secret CFO, September 20, on what good looks like with no AI in the picture: "Today, a three-day month-end close is absolutely world class. In a complex organization, getting there can take years of transformation work, flawless systems and processes."

Read them together. The human bar is three days after years of process work. The model bar is a fifth of expert tasks finished with three attempts. Any pitch for a continuous close has to clear both, and most of them are only arguing with the first.

What to do with the number

This is not an argument for waiting. It is an argument for scoping, and it produces one question you can ask of any finance workflow somebody wants to automate this quarter.

Given 21%, which of these workflows has a reviewer who would catch the other 79%, and what does that reviewer cost?

Where the answer is "a qualified person already checks this," you have a candidate: the model does the first pass, the human keeps the judgment, the saving is real and measurable. Where the answer is "nobody checks it, that was the point," you have a liability with a demo attached.

Note what the protocol itself assumes. Three attempts per task is generous. A firm that gets three shots at a close entry has somebody deciding which of the three to keep. That reviewer is the cost the benchmark does not price, and it is the cost every autonomy pitch leaves out.

Three things to hold loosely

The task list is not public, only the count and the rubric method, so nobody outside Ramp can audit what "everyday accounting workflows" covers.

Ramp sells finance automation, and a company in that position benefits from "this is hard and we are the ones doing it properly" exactly as much as a competitor benefits from the opposite. The number cuts against the obvious vendor interest, which is why it is citable, but it is still a vendor's number.

And nobody has run Mercury Books against this benchmark. Until somebody does, the contradiction stays a contradiction rather than a verdict.

What will not move: an open-weights model finished more real accounting tasks than either US frontier lab, and it still failed four out of five. That is the state of agentic accounting in September 2026, published by the people selling it.

Sources
1Ramp Labs, Ramp Accounting Bench, released 2026-09-17. First-party: "We partnered with accounting professionals to create 137 tasks and grading rubrics grounded in everyday accounting workflows. Even with three attempts, the best model achieved only 21% accuracy. Reliable agentic accounting remains an open problem."
2Per-model results from the companion post: Claude Fable 5.1 led the partial-credit average at 49.7%; GPT-6 Astra scored 48.7% "at roughly one third the cost"; GLM-5.3 (Z.ai) fully solved the most tasks at 21.2% with up to three attempts.
3Ian Tracey on X, 2026-09-18, on vertical benchmarks outside the training corpus.
4Immad Akhund on X, 2026-09-16, launching Mercury Books. Alex Rampell (a16z) the same day: "Your books can't be out of date if your bank = your books!"
5The Secret CFO, 2026-09-20, on the month-end close: "Today, a three-day month-end close is absolutely world class."
6The benchmark's task list is not published, only the count and the rubric method. Ramp sells finance automation. No source has run Mercury Books against this benchmark.
John Tan
John Tan

Founder and CEO of nativefirst.ai. Embeds with scaling founders and CEOs to ship Level-3 agents and AI workflows in production.