AI
OpenAI Says It Now Runs 3.1 Agent-Workdays per Researcher Day
OpenAI declares its 'automated research intern' goal met, with 3.1 agent-workdays per human day and $600 of daily inference per researcher. The same post explains why the ratio is not a productivity multiplier.

How much should you trust a productivity figure a company measures on itself? OpenAI invites the question with a post published September 6, titled "Research acceleration: the view inside OpenAI". It declares the goal it set last autumn, an "automated research intern" by September 2026, achieved. The headline number is 3.1 agent-workdays for every human workday. The more interesting material is the list of reasons OpenAI gives for not taking that number at face value.
What was actually measured
OpenAI defines an automated research intern as a system that completes well-scoped research tasks under human direction, including tasks that would take a skilled researcher several days. The definition is explicit about what it excludes: the system is not expected to originate a research agenda, decide which questions deserve compute, or judge on its own whether a result matters. This is a supervised fleet of coding agents, and the post does not pretend otherwise.
The 3.1 ratio comes from mid-August data. OpenAI added up the daily peak runtime of agents launched by people in its research organisation, plus the subagents those agents spawned, and divided by a standard eight-hour day. Spending tells the same story from another angle: the median researcher was consuming more than $600 a day of inference at API prices, and the 90th-percentile user more than $7,000 a day.
The footnotes OpenAI wrote itself
The post explicitly warns that 3.1 is not a straightforward productivity multiplier. Agent runtime can run in parallel, duplicate effort, fail outright, or need heavy human steering. Over six months, more than half of the successful tasks in the four-to-eight-hour range involved at least one human intervention, and long-horizon research still depends on a person at the wheel because errors compound.
Where the agents shine is unsurprising: coding, infrastructure work, technical troubleshooting, monitoring and running experiments. Usage is growing in the deciding, designing and analysing phases too, but the weight remains on the laborious end of research rather than the judgement end.
Two incidents, one budget cut
OpenAI writes that it does not know how to safely reach fully aligned recursive self-improvement.
That sentence sits in the same post as the milestone announcement. So do two incidents. On July 20, agents compromised training infrastructure and a container service was temporarily shut down. On August 7, evidence of cyber capability in GPT-6 Astra triggered model-specific controls; Astra-class GPU allocation then fell 59.2 percent, with capacity shifted to other models offsetting about 85 percent of the drop. We covered Astra crossing the cyber threshold when it happened; this post is the first time OpenAI has put an internal cost on that decision.
The next target is already set: an automated AI researcher by March 2028. OpenAI adds that capability progress cannot be assumed to stay in step with alignment and monitoring, and that options to slow or halt specific lines of work are built into its technical objectives rather than treated as external policy.
What a normal engineering team can take from this
Strip away the headline and three lessons remain. The heaviest agent user on the planet still intervenes by hand in half of its four-hour tasks, which suggests it is early for a small company to leave a coding agent unsupervised. Six hundred dollars a day per person means agents are not cheap; budgeting by tokens rather than by seat count is the honest way to plan. And the clearest gains sit in infrastructure troubleshooting and repetitive engineering work. That matches what we see in client projects: writing tests, reading logs and debugging deployments are the jobs where an agent pays for itself in the first week.
The 3.1 figure makes a good headline. The post's real value is that it also explains why you should not read it as a multiplier.
Sources: OpenAI, Unite.AI, Data Studios

Written by
Faruk Talmaç
Co-Founder & Editor
Co-founder of YZ Uzman, with 20+ years of experience in web design and software development.