Back to blog
AI StrategyOctober 20, 2026 4 min read

Same Score. Forty Times the Bill.

AP
Angelo Pallanca
Digital Transformation & AI Governance

Series The Harness · 4 of 5

TL;DR

When this series started, the argument for the harness was about capability: same model, different body, sixteen points of difference. Three studies published between June and September 2026 say the gap has moved. Pass rates between harnesses now sit within zero to eight points, while the cost of reaching them varies by up to forty times. Gartner has already priced the consequence at more than fivefold per agentic workflow through 2028. Cornwall faced the same measurement problem in 1811 and fixed it with one published number.


Fifteen months ago the case for the harness was a capability case. Same weights, two different bodies, sixteen points apart on a coding benchmark. That was post one of this series and it held up.

Three papers published between June and September 2026 say that is no longer the interesting part.


The scaffold barely moves the score now

In June, Naman Vats and Oleg Golev at Sentient Labs ran 300 trials: two models, Qwen 3.6 Plus and MiniMax M2.5, across three open-source harnesses, Goose, OpenCode and OpenHands-SDK, on a fifty-task stratified subset of Terminal-Bench Pro.

Pass rates clustered between 38 and 50 percent. Paired differences between harnesses on the same model: zero to eight points.

Then they counted tokens.

Cost per solved task ran from about 28.000 tokens to about 1,55 million. A 40x to 55x spread, for the same work, at the same success rate.

OpenCode produced roughly ten times more no-action turns than Goose. Idle loops. The agent thinking about thinking, metered.

Two models and fifty tasks is a small study, and open-source harnesses are not what most large companies actually run. Take the magnitude, not the decimal. The magnitude is the whole story.


The bill is an architecture problem

A July paper on enterprise orchestration held the task fixed across six models and varied only the harness. Tokens per task fell 38 percent, cost per task 41 percent, median latency 44 percent, with quality flat.

The authors give the failure mode a name: token maxing. Quality bought through steadily growing token intensity with falling marginal returns, masked by falling per-token prices and visible only in the invoice.

Late September, NVIDIA published SoL-Pi, a system that searches for better harnesses automatically. 152 directions, 535 executable environments, more than 3.000 runs. The efficient configuration used 49 percent fewer tokens and kept 93,7 percent of the score.

Sit with that. One of the largest hardware companies on earth spent thousands of compute runs optimising the body, not the brain.


Gartner has already priced it

On 17 August 2026 Gartner predicted that inference cost per agentic workflow will rise more than fivefold through 2028. Their analyst Will Sommer described it as better unit economics escalating the overall cost of AI without a clear pathway to commensurate value.

Per-token prices fall. Bills rise. Both true. The harness is where the two meet.


Cornwall solved this in 1811

Cornish mine engines burned coal to pump water, and nobody could compare them honestly, because horsepower says nothing about fuel. From 1811 a monthly report began publishing each enginès duty: millions of pounds of water lifted one foot per bushel of coal. Not power. Work per unit of fuel.

Average duty in Cornwall went from roughly twenty millions in the 1810s to over eighty by the late 1820s. The engines did not receive new physics. They received a public number, monthly, with everyonès name on it.

We still rank agents by pass rate the way Cornwall once ranked engines by horsepower. Some leaderboards, Artificial Analysis among them, now publish cost and tokens per task next to the score. The instrument exists. It is not yet the unit people quote.


Why this matters for your business

If you are picking an agent platform on benchmark score, you are picking on the metric with a zero to eight point spread and ignoring the one with a forty times spread.

Ask your vendor for tokens per completed task, not per call. Ask what share of turns take no action at all. Then run the pilot on your own work and read the invoice, not the leaderboard.

The brain is rented. The body is what you are billed for.

Want to discuss this further?

Request a proposal