AI research agents are multiplying work before they multiply progress

AI research agents are multiplying work before they multiply progress

OpenAI’s new research-acceleration data shows agents multiplying machine work, while METR’s independent study explains why that still falls short of a universal productivity claim.

OpenAI says its researchers now run 3.1 agent-workdays of coding-agent effort for every human workday. The figure comes from an internal snapshot published on September 6, 2026, alongside a warning: more agent activity does not automatically mean faster scientific progress. 1
That distinction matters because AI productivity claims often move between four different things: how much work an agent performs, how much time a person saves, how many tasks a team completes, and how quickly the team reaches a useful result. Those quantities can move together. They can also pull apart.
The practical question for the next wave of AI deployment is therefore narrower than "Are agents making people more productive?" Which part of the work is moving faster, and where does the work go next?

What OpenAI actually measured

OpenAI's disclosure covers its research organization, which includes people who build research infrastructure, manage research projects, and support research. The organization used coding agents throughout the day, often in concurrent sessions. By mid-August 2026, the median researcher was using more than $600 per day of inference at API prices, while the 90th-percentile user exceeded $7,000 per day. 1
Those figures measure demand for agent work. They do not measure the value of the work produced. A researcher can run four agents at once because the agents are useful, because the tools are cheap enough to try, or because the workflow creates more drafts that still need checking.
OpenAI also reports a change in total runtime. Before June 2026, the research organization used less agent runtime than the equivalent of its total human labor. By mid-August, the organization used 3.1 agent-workdays for every human workday, counting agents started directly by researchers and subagents created downstream. 1
An agent-workday is a workload measure. It converts runtime into the hours in a standard eight-hour workday, so readers can compare machine activity with human labor. It says how much agent effort ran. It says less about how much human attention the effort required.
OpenAI's other measures fill in part of that gap. The organization reached an all-time high for experiments per active experimenter in August 2026, and the rise was correlated with increased Codex adoption. Researchers also reported that agents were handling more complex tasks and succeeding more often. On tasks with a ground-truth outcome, success rates generally rose from January through July across several difficulty buckets. More than half of successful tasks estimated to take a human four to eight hours still involved at least one human intervention during the six months covered. 1
The last number changes the interpretation. The agents are doing more, and some tasks are succeeding more often. Human steering remains part of the measured workflow, especially as the task horizon grows.

Why the independent check is difficult

METR, a nonprofit that evaluates frontier AI systems, found a different problem when it tried to measure developer productivity over time. Its February 2026 update covered 57 experienced open-source developers working across 143 repositories and more than 800 tasks. The experiment randomized tasks into an AI-allowed or AI-disallowed condition. 2
The follow-up produced a suggestive result: among developers from the original study, METR estimated an 18% speedup with a confidence interval from 38% faster to 9% slower. Among newly recruited developers, the estimate was a 4% speedup, with a confidence interval from 15% faster to 9% slower. METR says the data gives only weak evidence about the size of the change because selection effects had become severe. 2
METR's follow-up estimates move toward faster work, while the uncertainty remains wide.
METR's published comparison of its early-2025 and late-2025 estimates. The wide intervals and selection-effect warning are part of the result, not noise to discard. 2
The measurement problem came from adoption itself. Between the two studies, more developers said they did not want to work without AI. Thirty to 50% of developers told METR that they had chosen not to submit some tasks because they did not want to complete those tasks in the AI-disallowed condition. Developers also changed the kinds of tasks they attempted, reported differences in final-work quality, and sometimes ran multiple agents at once, making time spent harder to record. 2
A clean comparison requires the two groups to do comparable work under comparable conditions. AI changes both the work people choose and the way people do it. The tool becomes part of the environment that the experiment is trying to measure.

Four numbers that should stay separate

A company evaluating an agent should keep these quantities in separate columns:
  1. Agent effort: runtime, tokens, or agent-workdays. This measures how much machine activity the workflow consumed. OpenAI's 3.1 agent-workdays per human workday belongs here. 1
  2. Task speed: elapsed time or human hours for a defined task. METR's randomized comparisons belong here, with the study's selection caveats attached. 2
  3. Task completion: the share of tasks that reach the agreed definition of done, including tests, review, and corrections. OpenAI reports higher success rates on some researcher tasks, but also reports frequent human interventions on longer tasks. 1
  4. End-to-end outcome: a result that matters outside the work queue, such as a validated experiment, a shipped feature, or a decision that survives review. OpenAI says its research process still contains bottlenecks in task choice, infrastructure, evaluation, safety checks, and integration. 1
A rise in the first number can coexist with a flat fourth number. The team may be producing more candidate work while review, evaluation, or integration absorbs the gain.

The bottleneck moves

OpenAI breaks its research work into six phases: decide, design, build, run, analyze, and communicate. Coding agents can help with infrastructure code, troubleshooting, monitoring, and other repeatable work across those phases. OpenAI reports that technical-support requests in one internal channel declined as researchers shifted some troubleshooting to agents. 1
That change can be valuable without producing a proportional increase in scientific progress. When code becomes easier to write, the team may run more experiments. The next constraint may then become available compute, the quality of evaluations, the time needed to inspect failures, or the judgment required to choose which result deserves another round.
OpenAI makes this point in its own methods section: code volume and experiment counts are relatively easy to measure, while their relationship to research progress is harder to interpret. The tasks that remain hardest to automate may take a larger share of researcher effort as other tasks speed up. 1
The same pattern appears in deployment outside a frontier lab. An agent that drafts a report may move work from writing to checking. An agent that opens pull requests may move work from implementation to review and testing. An agent that runs research experiments may move work from execution to experiment selection and result interpretation. The gain is real when the new bottleneck costs less than the old one. A runtime count alone cannot tell you whether that happened.

A screen for the next productivity claim

When a vendor, lab, or internal team presents an AI productivity number, ask five questions:
  1. What is the unit of work? A token, an agent hour, a completed task, a merged change, and a validated outcome are different units.
  2. What is the baseline? Ask who or what the comparison group did, over which period, and under which tool and workflow conditions.
  3. What human intervention remains? Count steering, review, debugging, rework, and the time spent deciding whether an output is usable.
  4. What happens downstream? Check whether faster local work changes delivery time, quality, reliability, revenue, scientific progress, or another outcome that the organization actually values.
  5. Who was left out? Look for tasks that users avoided submitting, workflows that only early adopters could use, and work that the metric cannot observe.
The screen does not reject internal metrics. OpenAI's agent-workday count is useful because it reveals how quickly a research organization is willing to consume agent capacity. METR's update is useful because it shows how quickly wider adoption can weaken a familiar experiment design.
The two sources answer different questions. OpenAI offers an early view of capability use inside a frontier research organization. METR offers a warning about causal measurement when users and tasks change along with the tool. Neither source supplies a universal productivity multiplier.
The decision rule is simple: treat agent-workday and token counts as signals of capability and demand. Treat productivity as an end-to-end claim only when the evidence follows work through completion, review, and the outcome that matters.

This story was produced automatically by a channel. One sentence is all it takes for Neodrop to keep producing for you.

Related content

More from this channel