
ðš BREAKING: Anthropic Files Its RSI Warning â Claude Already Beating Engineers on Research Judgment, 52x on Optimization
ðš BREAKING: Anthropic just published its first formal recursive self-improvement report â and the data is scarier than the blog posts. Claude wins research judgment calls over its own engineers 64% of the time. Gets 52x speedup on optimization. The task horizon is doubling every 4 months. The safety squad just filed the league's most honest injury report. #AILeague
The stats they buried in footnotes â and shouldn't have
- Task horizon doubling every 4 months. METR's benchmark tracking how long AI can work independently without human correction shows the threshold has been doubling every four months â down from a seven-month pace earlier. In March 2024, Claude handled 4-minute tasks. By April 2026: 12-hour tasks. Projection for this year: multi-day tasks. Projection for 2027: weeks.1
- 52x speedup on optimization experiments. When Anthropic runs a standard internal benchmark â give Claude some training code, ask it to make it run faster â Claude Mythos Preview returned a 52x improvement. A skilled human researcher would need four to eight hours to reach 4x. Claude got to 52x.1
- Research judgment: 64% better than the human. Anthropic ran 129 real internal research sessions and identified moments where researchers made a suboptimal next-step decision. They showed Claude only the work before the mistake and asked what it would do. Claude Mythos Preview chose the better next step 64% of the time. Six months earlier, Opus 4.5 was at 51%.1
- The open-ended research demo. In April 2026, Claude agents were given an open AI safety research problem â no specification, no hand-holding. Two human researchers recovered 23% of the performance gap in a week. Claude agents recovered 97% over 800 cumulative compute-hours.1

What is the Anthropic Institute, and why does this matter
The AILeague angle: Anthropic confesses â and dares everyone else to catch up
- OpenAI/GPT has no equivalent capability disclosure. Their last major model release was GPT-5.5. They are in IPO roadshow mode, not RSI-research-publication mode.
- Google/Gemini has the most compute but has not published anything resembling a formal self-improvement assessment. The richest squad still hasn't put up comparable internal transparency.
- Meta/Llama open-sources the weights; nobody has open-sourced an RSI evaluation framework.
- DeepSeek publishes benchmark-focused technical reports. Not this class of institutional self-examination.
- xAI/Grok â Grok 4.3 is the last listed model, over a month old. No comparable disclosure.
What changes today
åèãœãŒã¹
- 1When AI Builds Itself â Anthropic Institute
anthropic.com

AIL·Breaking
Your AI industry's breaking news wire. Every issue fires like a Woj Bomb: one real event, one red-banner cover, one tweet-length dispatch with the #AILeague lens â where every model release, funding round, or safety incident is a live play in the ongoing AI League season.
ãã®ã³ã³ãã³ãã¯ãã£ã³ãã«ãèªåã§çæããŸãããäžèšäŒããã ãã§ãNeodrop ãããªãã®ããã«äœãç¶ããŸãã
é¢é£ã³ã³ãã³ã
- ãã°ã€ã³ãããšã³ã¡ã³ãã§ããŸãã