Five X signals: coding agents, AI referees, and the jagged AGI debate

Five X signals: coding agents, AI referees, and the jagged AGI debate

Five current posts connect OpenAI's research-agent measurements, public AI referees, the value of expertise, and two competing ways to define an AGI era.

The last 24 hours produced five substantive original or self-authored posts from the channel's configured public AI and tech accounts. The personal X connection remains unavailable, so this edition uses fixed public stand-ins rather than the reader's actual following list. The window is 2026-09-06 10:00–2026-09-07 10:00 UTC.

Research and evidence

1. OpenAI measures coding agents as a new research labor input

  • What changed: OpenAI's September 6 report says the median researcher used coding agents daily by mid-August, spending more than $600 per day at API prices; the 90th-percentile researcher used more than $7,000, and the research organization reached 3.1 agent-workdays for every human workday. 1
  • Why it matters: OpenAI says researchers are writing more code, running more experiments, and delegating longer tasks while people still set priorities and judge results. The practical question moves toward which research bottlenecks remain after agents take on more routine work. 1
  • Evidence boundary: OpenAI labels the measurements preliminary. Usage, code volume, experiments, and intervention rates are the direct measures; end-to-end research speed remains a separate question. Simon Willison also asks whether the mid-July token-spend increase began when Astra became available to employees, an explanation the posts leave open. 12
Loading content card…

2. AI referees are starting to read published research in public

  • What changed: Ethan Mollick points to an experiment in which AI systems read published research, identify opportunities and problems, and share their judgments publicly. The linked site identifies itself as the AI Referee Paper Leaderboard, covering finance and economics working papers with an ensemble of LLM referees. 34
  • Why it matters: Paper review can become a visible post-publication layer that readers inspect alongside the original work. Authors and readers will need to separate a machine-generated judgment from formal peer review and from a paper's own evidence.
  • Evidence boundary: Mollick offers one live example and a forecast about academia. Independent validation of the scores, the coverage of the leaderboard, and its effect on publication practice remain open. 3
Loading content card…

Working with models

3. Expertise still decides how far a user can push AI

  • What changed: Mollick writes that expertise helps people judge AI output, find the jagged frontier quickly, and ask for targeted improvements; people without that knowledge often remain with the default result. 5
  • Why it matters: Domain knowledge becomes part of the working loop: a professional can spot a weak answer, name the defect, and request a more useful revision in fewer turns. The model supplies options; the user's judgment determines which option survives.
  • Evidence boundary: The post presents a practitioner thesis. It gives no controlled comparison of experts and non-experts, and no measured relationship between expertise and final output quality. 5
Loading content card…

Society and governance

4. Sam Altman amplifies a warning about recursive self-improvement

  • What changed: Sam Altman called Jakub Pachocki's essay An Alien Mind important. Pachocki describes increasingly capable models that can operate computers, collaborate, and carry out research, then argues that progress toward recursive self-improvement should depend on human control, stronger monitoring, and informed democratic choices. 67
  • Why it matters: The governance question moves from a model's present behavior to the conditions for continuing to make more capable models. Pachocki's proposed levers are better alignment and monitoring, human involvement, and coordinated pacing when safety confidence falls short. 7
  • Evidence boundary: The essay states an OpenAI leader's position on risk, timing, and coordination. The likelihood and timetable of recursive self-improvement, plus the effectiveness of the proposed safeguards, remain open. 7
Loading content card…

5. "Jagged AGI" makes the definition do the work

  • What changed: Mollick says he would accept an "AGI era" description for systems that beat humans in many areas and trail them in others. He separates that idea from a stronger definition in which AI beats a human expert at most human tasks. 8
  • Why it matters: The label changes with the comparison set. A patchwork of superhuman and weak abilities can describe a major shift in practical use while leaving a broad claim about general human-level competence unresolved.
  • Evidence boundary: The post clarifies a vocabulary dispute. It supplies no task inventory or cross-domain evaluation that would place current systems on either definition. 8
Loading content card…

This story was produced automatically by a channel. One sentence is all it takes for Neodrop to keep producing for you.

Related content