AI
METR's New Metric Prices AI Agents Against Human Experts
METR's expenditure horizon finds the budget at which a human expert becomes cheaper than an AI agent. On a hard optimization task, tested agents ran out of value between $0 and $3,300.

At what budget does an AI agent stop being the cheap option? Most vendor pitches never reach that question, because most benchmarks answer a different one: can the model do the task, yes or no. METR has published a metric built specifically to answer the cost version.
The research organization calls it the expenditure horizon. The idea is straightforward. Measure how much money an AI agent needs to reach a given improvement on a task, measure how much a human expert needs to reach the same improvement, and find the budget where the two curves cross. Below that crossing point, the agent is the better buy. Above it, hiring the human is.
How the expenditure horizon works
Two design choices make the metric more useful than a pass-fail score. It produces a continuous value, improvement per dollar, instead of a verdict. And it converts everything into one currency: token costs, experiment compute and human labor all land in the same column, which is the only way a comparison like this survives contact with a real budget.
The inputs it needs are the ones a finance team already tracks. Compute and operating cost on the AI side, hours and hourly rate on the human side, and a quantifiable measure of progress on the task itself. That last requirement is the constraint: the method needs work whose improvement can be measured on a number line.
The numbers from the NanoGPT test
METR ran the metric on the NanoGPT speedrun, an optimization task with a long public history. Across 82 improvement steps, contributors have cut training time from 45 minutes to under two minutes, a 33x speedup since May 2024. Human effort on the project totals roughly $250,000, working out to about 16 hours per percentage point of speedup at an assumed $150 per hour, or around $2,500 per one-percent gain.
Six models were then given $10,000 each. GPT-5 and Opus 4.1 produced no real progress. GPT-5.5 managed about 1% improvement and Opus 4.8 about 1.5%, both with estimated expenditure horizons in the low four figures. Across the tested models the horizons landed between $0 and roughly $3,300.
METR's own summary of the result is blunt: autonomous optimization has barely moved the needle on NanoGPT progress so far.
Bringing this back to a purchase decision
Read carelessly, this looks like evidence that agents do not work. Read properly, it says something narrower and more useful: on open-ended research optimization, where the next improvement requires genuine insight rather than execution, agents currently run out of value at a few thousand dollars of spend. That is one task family, deliberately chosen because it is hard.
The practical value is in the shape of the question, not this particular answer. Before committing to an agent deployment, the useful thing to establish is not whether the agent can do the task but at what budget it stops being worth it, and whether the work you have in mind resembles NanoGPT-style open-ended optimization or the far more common category of well-specified, repetitive execution where agents currently earn their keep.
Most of the automation work that pays for itself sits firmly in the second category. Knowing which category a project belongs to before the budget is approved is the difference between an automation program and an experiment with a rollout plan attached.
Sources: METR, The Decoder

Written by
Faruk Talmaç
Co-Founder & Editor
Co-founder of YZ Uzman, with 20+ years of experience in web design and software development.