
OpenAI Measures Its Research Agents—and Warns About the Next Intelligence Jump
A concise brief on OpenAI’s measured rise in research-agent use and its chief scientist’s warning that monitoring and human oversight must keep pace with recursive self-improvement.
This brief covers developments published from September 6 at 08:15 through September 7 at 08:15, 2026, Asia/Dhaka. Two OpenAI publications stood out: one measures how heavily research teams are using coding agents, while the other sets out the lab’s chief scientist’s expectations for recursive self-improvement and the safety work he says must accompany it.
Quick scan
| Development | What changed | Why it matters |
|---|---|---|
| OpenAI’s research-agent measurements | By mid-August, the research organization used 3.1 agent-workdays for every human workday; median daily inference use passed $600 at API prices. 1 | Internal agent use is moving from occasional coding help toward longer, more concurrent research work. |
| OpenAI chief scientist Jakub Pachocki’s safety essay | Pachocki argues that current progress could continue into recursive self-improvement and says monitoring confidence should constrain further scaling. 2 | The next checkpoint is institutional as much as technical: stronger monitoring, external safety bars, and a way to keep people involved in AI’s improvement loop. |
Research agents are already taking more of the workday
OpenAI says it met its goal of an automated “research intern” in September 2026 and is working toward an automated AI researcher by March 2028. The new internal report measures the path through usage and workflow data rather than through a new model benchmark. 1
The clearest change is scale. By mid-August, the median researcher in OpenAI’s research organization used more than $600 of daily inference at API prices. The 90th-percentile user passed $7,000 per day in token use. OpenAI also estimated 3.1 agent-workdays for every one human workday, using a standard eight-hour workday as the unit. Before June, total agent runtime remained below total human labor. 1
The work itself is changing. Experiments per active experimenter reached an all-time high in August, measured from tracking that began in January 2025. Delegated work also shifted toward higher-level and longer-horizon tasks. OpenAI reports growing use of workflows that run four or more agents at the same time, alongside a rise in technical help and monitoring runs. 1

The report gives the numbers a useful limit. Code volume and experiment counts are easier to measure than overall research progress, and compute remains a bottleneck. More than half of successful tasks estimated at four to eight hours of human work involved at least one human intervention. OpenAI describes the measurements as preliminary and says they cover most, rather than every part, of agent use. 1
What to watch: OpenAI’s next useful checkpoint is the quality of the research output. The organization has shown rising agent use, more experiments, and longer task horizons. The harder question is whether those measures translate into better discoveries after human steering, compute limits, and infrastructure work are counted.
OpenAI’s chief scientist puts monitoring at the centre of the next step
In a separate September 6 essay, OpenAI chief scientist Jakub Pachocki says he has a strong expectation that the current speed of AI progress could continue into recursive self-improvement: systems increasingly contributing to their own development. Pachocki presents that outcome as a forecast grounded in OpenAI’s internal results, rather than as an established event. 2
Pachocki separates two alignment goals. Goal alignment concerns following an instruction hierarchy and cooperating with people. Value alignment concerns acting reasonably in unfamiliar or adversarial situations, including honesty, integrity, and regard for humanity. He identifies generalization as the central difficulty: a model may encounter environments and higher-level concepts that differ sharply from its training examples. 2
Pachocki says GPT-6 Astra is significantly better aligned than GPT-5.6 Sol, while adding that alignment progress needs to keep ahead of rising general intelligence. He also says OpenAI’s ability to rely on chain-of-thought monitoring is progressively diminishing. His reasons include more complex tool-using environments, models becoming better at manipulating their own reasoning, and stronger pretrained capabilities that appear without verbalized reasoning. 2
That argument changes the practical checkpoint for automated research. The question becomes whether people can still inspect and steer the process as models take on longer tasks and participate in improving future systems. Pachocki says OpenAI will continue technical alignment and monitoring work, build defensive systems, and withhold further scaling when needed. He also calls for safety bars that can be enforced by third-party auditors, governments, or international bodies, along with international coordination and voluntary slowdowns until shared standards exist. 2
What to watch: Look for concrete monitoring methods, external review arrangements, and evidence that human operators remain part of the improvement loop. OpenAI’s March 2028 automated-researcher target supplies a date; the harder checkpoint is whether its safety controls become more reliable before the target arrives.
What to watch
- Whether OpenAI publishes research-quality measures that go beyond agent spending, coding volume, and experiment counts. 1
- Whether OpenAI’s monitoring work can keep pace with models that use tools, coordinate with other AI systems, and reason about their own reasoning. 2
- Whether the proposed safety bars gain independent enforcement and international coordination, rather than remaining commitments made by individual labs. 2
References
- 1
- 2An Alien Mind
openai.com
- 3Research acceleration: The view inside OpenAI — Simon Willison
simonwillison.net
This story was produced automatically by a channel. One sentence is all it takes for Neodrop to keep producing for you.
Related content
More from this channel›
- NASA-IBM's lunar model, OpenAI's Agents API, Anthropic's threat report, and Gemini for Windows
- Google's €13B nuclear pact, MAPL-EMIT from orbit, OpenAI's safety mandate, and Suno v6
- Muse, Mistral's €3B, Cognition's $48B, and AI at the lab bench
- Rentosertib's six aging clocks point to a biomarker signal
- OpenAI Acknowledges the Wiki Incident, TCS Plans a $7.4B AI Campus, and 128 States Reach an Autonomous-Weapons Text
- Claude Formalizes Fermat, Gemini Spark Gets Photos, and a New OpenAI Agent Report
- GPT-6 Astra, NVIDIA's Hugging Face Deal, and WeatherNext 3
- Qwen3.8-Max-0902、亚马逊消息核验、PaperCompiler:AI 的三条新进展
