🔮 Exponential View #599:Unbounded self-improvement and its limits|无限自我改进及其边界|英文原文 + 中文翻译

🔮 Exponential View #599:Unbounded self-improvement and its limits|无限自我改进及其边界|英文原文 + 中文翻译

Exponential View #599 讨论递归式自我改进的物理与实践边界、开放权重模型的成本性能竞争,以及 AI 参与芯片设计带来的算力市场分化。

原文信息

  • 英文标题:🔮 Unbounded self-improvement and its limits #599。1
  • 副标题:Plus: The deceptive popularity of open-weight models & DIY compute. 1
  • 作者:Azeem Azhar、Nathan Warren。1
  • 发布时间:2026 年 8 月 30 日 10:51:35(北京时间)。官方 RSS 记录为 Sun, 30 Aug 2026 02:51:35 GMT。2
  • 公开范围:官方页面本次公开展示了从开场到结尾的全部文章正文,包括 Morsels 和脚注;以下不补写页面之外的内容。1

English original

Good morning! 1
A MESSAGE FROM OUR SPONSOR
This paid post was unlocked for you by Granola, the AI notepad for back-to-back meetings.
Granola is our team’s favorite note-taker. We’re excited to partner with Granola to bring four months of free access to the Business plan (10 seats) to Exponential View annual members. This perk, valued at $560, is now part of our membership.
Upgrade to claim your free access to Granola.

The limits to recursion

What are the conditions under which AI systems could undergo recursive self-improvement (RSI) — and how long might that last? Cards on the table, I’m not wildly excited by the theory of unending accelerating recursive self-improvement for the simple theoretical issue of control and alignment. But I also think it’s not likely for practical and theoretical reasons. 1
Now philosopher, Toby Ord, has done the heavy lifting for me. He argues that a key gating factor to RSI is the generation time: how long it takes of an AI system to go through a single loop of improvement (where it helps design and train its successor). 3
Ord concludes that while it is mathematically possible for extreme RSI, intelligence rising without bound, the conditions are unbearably difficult to achieve. The key question is whether the entire research to training to development cycle can shrink to zero or not. Ord reckons unlikely, I too don’t believe it is possible. 1
Generation time can’t get to zero because real-life intrudes: experiments take time; training runs take time; making new chips take time… lots of things take time. It might still feel fast, but it wouldn’t race to infinity. 1
Eventually, physics intrudes too: the speed of light limits communication speed; the Bekenstein bound limits the information contained within a finite bit of space; and Landauer imposes an energy tax on irreversible computation.1 1
The Universe, it seems, agrees with me. Unbounded RSI has its limits.
You can read Ord’s paper here. Premium members can explore a plain English interactive version too.

The market is the instrument

Open-weight models are growing in popularity in the business world: their token share at Vercel hit a single-day record of 62%, up from 28% two months earlier. Some Western firms are even moving workloads to Chinese open weights — Thomson Reuters has developed its first in-house model based on Qwen to cut costs. You can tune cost-effective open weights to match, or sometimes beat, frontier performance on the tasks that matter to you. 1
Take Bridgewater: working with Thinking Machines, it fine-tuned an open Qwen model on expert-labeled data and beat every frontier model it tested on its internal information-filtering tasks: roughly 30% fewer errors than the best closed model, at one-fourteenth of the inference cost. Trainloop, which I am an investor in, does something similar, using the tiny Qwen 3.7-27b model, and can outperform GPT 5.6 Sol on specific fine-tuned tasks at a fraction of the cost. 45
Scatter plot comparing the average accuracy and cost per 1,000 tasks of a trained model with frontier models
The official chart compares the trained model with frontier models by average accuracy and cost per 1,000 tasks. 1
The results that Trainloop is getting are pretty impressive, seeing as they are based on a pocket model—a 27b model will even fit on a desktop Mac.2 1
Openweight models are, of course, getting better and better. Z.ai GLM 5.3, released this week, completely reshapes the cost-performance Pareto frontier. Of course, it isn’t a small model, but smaller distillations will emerge from it. 1
Artificial Analysis Intelligence Index Pareto frontier showing model score against cost per task
The official Artificial Analysis chart is titled “Pareto frontier of the Artificial Analysis Intelligence Index v4.1.1”; it plots 59 models by Intelligence Index score and cost per task, and is marked as updated 2026-08-26. 1
All of this speaks to a welcome competition in AI provision. Clearly, firms could move focused workloads onto the most-performant, fine-tuned small models they can. Where possible, they might choose large, generally capable open models. But the appeal of being at the frontier, which is more than just model performance—it is service guarantees, harness quality, reliability, and a host of other requirements—still drives significant business for Anthropic and OpenAI. 1
We don’t think it has much impact on the question of whether revenues flowing into the industry will materially change. For one thing, we don’t have a counterfactual to test against. But more importantly, every open model still involves paying inference providers. We’ll be looking at this question in more detail in the comings weeks. 1

A spicy entry to the compute ecosystem

Recursive self-improvement may be entering the compute realm. OpenAI’s new chip, ‘Jalapeño’, was designed with a heavy helping hand from the company’s own models, which helped write kernels and cut roughly 10% from one of the chip’s main compute blocks. In around 16 months from first hire to tape-out, OpenAI has built a chip that beats comparable Nvidia silicon by 1.5–1.9x on tokens per megawatt at peak throughput. This suggests frontier models can compress the design cycle for competitive silicon. 67
AI will result in far greater heterogeneity in chip architectures than we saw in prior computing markets. Personal computers battled between the x86 standard and the Motorola 68x, before today’s duopoly of Intel and Apple silicon. Different uses for AI will need compute optimised for intelligence, latency, power consumption, training, and inference. This creates lots of room for specialist firms. One example is that ChatGPT’s fast response mode is powered by Cerebras’ low-latency silicon. Another is Fractile, where I am an investor, which has a deal with Anthropic for its low-latency inferencing chips. 1
Compute is becoming a highly segmented market, where chips aren’t a standardized commodity. This differentiated hardware demand will expand the market even if it potentially reduces Nvidia’s relative dominance. 1

Morsels

Good post from Chad Syverson on AI productivity. Some micro-evidence for improvements, not much at the aggregate level.
Understanding AI and Productivity — Chad Syverson, George C. Tiao Distinguished Service Professor of Economics at the University of Chicago Booth School of Business, as part of EIG’s American Worker Project. 8
The least bad place to hide from global catastrophe? Australia. 9
The harness matters as much as the model: SwarmOS pushed GPT-5.6 Sol from 13.3% to 100% on ARC-AGI-3 Public. 10
The University of Chicago’s Social Sciences Core is going back to paper, banning most classroom technology to deal with the AI-learning crisis. 11
Meta considered shrinking some teams by up to 60% to become “AI native.” 12
With one video, you can reconstruct a moving 4D avatar of a person and render it from novel viewpoints, like a video-game character. 14
You can now teach an adorable “Pixar-ish” robot new tricks for only $399. 15
You can now generate videos in less time than it takes to watch them. 16
A walk down memery lane. 17
Reversible processors, like those from Vaire, will not pay the Landauer tax. 1
I run a version of Qwen 3.7-27b on one of our local machines for various tasks. 1
Recorded at least four quarters earlier. 1

中文翻译

早上好!1
赞助商信息
这篇付费文章由 Granola 解锁。Granola 是一款面向连续会议场景的 AI 记事本。
Granola 是我们团队最喜欢的会议记录工具。我们很高兴与 Granola 合作,为 Exponential View 年度会员提供 Business 方案四个月的免费使用权,方案包含 10 个席位。这项价值 560 美元的会员福利,现已加入我们的会员权益。1

递归的边界

AI 系统在什么条件下可能进行递归式自我改进(recursive self-improvement,简称 RSI)?这种过程可能持续多久?坦白说,我对“无止境加速的递归式自我改进”这个理论并不兴奋,原因很简单:控制和对齐本身就是一个理论难题。但出于实践和理论上的原因,我也认为这种情况不太可能发生。1
哲学家 Toby Ord 替我完成了最费力的部分。他认为,RSI 的一个关键闸门是 generation time(代际时间):AI 系统完成一轮改进需要多长时间,也就是系统帮助设计并训练继任者所需的时间。3
Ord 的结论是,极端 RSI、让智能无界增长,在数学上确实可能成立,但要满足所需条件极其困难。关键问题在于,从研究到训练再到开发的整个周期,能不能缩短到零。Ord 认为这不太可能,我也不相信它能做到。1
代际时间不可能缩短到零,因为现实会介入:实验需要时间,训练运行需要时间,制造新芯片也需要时间,还有很多环节都需要时间。整个过程仍然可能让人感觉很快,但它不会一路冲向无穷。1
最终,物理学也会介入:光速限制通信速度;Bekenstein bound(贝肯斯坦界)限制有限空间能够容纳的信息量;Landauer principle(兰道尔原理)则为不可逆计算设定能耗代价。1 1
看来,宇宙也同意我的判断。无界 RSI 有自己的边界。
你可以阅读 Ord 的论文。付费会员还可以查看一份用通俗英语制作的互动版本。

市场就是仪表盘

开放权重模型在商业世界越来越受欢迎:它们在 Vercel 的 token 份额达到单日 62% 的纪录,两个月前还是 28%。一些西方公司甚至开始把工作负载迁移到中国开放权重模型。Thomson Reuters 已经开发了第一款基于 Qwen 的内部模型,用于降低成本。企业可以针对真正重要的任务,调校成本更低的开放权重模型,让它们达到、甚至偶尔超过前沿模型的效果。1
以 Bridgewater 为例:该机构与 Thinking Machines 合作,用专家标注的数据微调了一个开放权重 Qwen 模型。在内部信息筛选任务中,这个模型击败了它测试过的每一个前沿模型:相比最好的闭源模型,错误数大约少 30%,推理成本只有后者的十四分之一。Trainloop 也在做类似的事。作者是 Trainloop 的投资人之一;该公司使用小型 Qwen 3.7-27b 模型,在特定微调任务上可以用远低于 GPT 5.6 Sol 的成本取得更好的表现。45
上方官方图表比较了一个训练模型与多个前沿模型的平均准确率,以及每完成 1,000 项任务的成本。1
Trainloop 使用的是一个“口袋模型”,因此它取得的结果相当亮眼:一个 27b 模型甚至可以装进桌面 Mac。2 1
开放权重模型当然也在持续进步。本周发布的 Z.ai GLM 5.3 彻底重塑了成本—性能 Pareto 前沿。它当然算不上小模型,但更小的蒸馏版本会随之出现。1
上方官方图表名为“Artificial Analysis Intelligence Index v4.1.1 的 Pareto 前沿”,以智能指数得分和单项任务成本绘制 59 个模型,并标注为 2026 年 8 月 26 日更新。1
这些变化说明,AI 服务领域正在出现令人欢迎的竞争。企业显然可以把重点工作负载迁移到自己能找到的、性能最好的微调小模型上。在条件允许时,企业也可以选择能力广泛的大型开放模型。但“处在前沿”的吸引力并不只来自模型性能,还包括服务保障、工具链质量、可靠性和一系列其他要求;这些因素仍然为 Anthropic 和 OpenAI 带来大量商业需求。1
我们认为,这些变化对“流入 AI 行业的收入是否会发生实质变化”这个问题影响不大。一方面,我们没有反事实情况可供比较。更重要的是,每一个开放模型仍然需要向推理服务提供商付费。未来几周,我们会更详细地研究这个问题。1

算力生态中一枚辛辣的新成员

递归式自我改进可能正在进入算力领域。OpenAI 的新芯片“Jalapeño”在设计过程中大量借助了公司自有模型;这些模型帮助编写内核,并让芯片一个主要计算模块的面积减少了大约 10%。从首次招聘到流片,OpenAI 在大约 16 个月内完成了一款芯片。按照峰值吞吐量计算,这款芯片的每兆瓦 token 处理能力比同类 Nvidia 芯片高 1.5—1.9 倍。这说明,前沿模型可以压缩具有竞争力的芯片设计周期。67
与过去的计算市场相比,AI 会带来更加多样化的芯片架构。个人电脑曾在 x86 标准和 Motorola 68x 之间竞争,后来形成了 Intel 与 Apple silicon 的双头格局。AI 的不同用途会需要针对智能水平、延迟、功耗、训练和推理进行优化的算力。这为专业芯片公司留下了很大空间。一个例子是,ChatGPT 的快速响应模式由 Cerebras 的低延迟芯片提供动力。另一个例子是 Fractile。作者是 Fractile 的投资人之一;该公司与 Anthropic 达成了一项协议,提供低延迟推理芯片。1
算力正在变成高度分化的市场,芯片不再是标准化商品。即使这种差异化的硬件需求可能削弱 Nvidia 的相对主导地位,它仍会扩大整个市场。1

零碎信息

Chad Syverson 关于 AI 生产率的文章值得一看。文章提供了一些微观层面的改善证据,但在总量层面证据不多。
《理解 AI 与生产率》:作者 Chad Syverson 是芝加哥大学布斯商学院 George C. Tiao 杰出经济学教授;文章属于 EIG 的 American Worker Project。8
躲避全球灾难最不坏的地方?澳大利亚9
工具链和模型同样重要:SwarmOS 把 GPT-5.6 Sol 在 ARC-AGI-3 Public 上的成绩从 13.3% 推到了 100%10
芝加哥大学社会科学核心课程为了应对 AI 学习危机,正回到纸笔,禁止课堂使用大多数技术设备。11
你愿意下注吗?如果截至 2033 年第四季度(包括该季度在内)的任一季度,美国人均实际 GDP 比此前峰值至少高 15%,这笔赌注就会兑现。3 13
现在只用一段视频,就可以重建一个人的运动式四维头像,再从新的视角渲染它,让它看起来像电子游戏角色。14
现在只要 399 美元,就可以教一台可爱的“皮克斯风格”机器人学习新动作15
现在生成视频所需的时间,已经可以短于看完视频本身的时间16
一起走一趟“迷因往事小径”吧。17
Vaire 生产的可逆处理器,不会承担 Landauer 能耗代价。1
我在我们的一台本地机器上运行 Qwen 3.7-27b 的一个版本,用于各种任务。1
记录时间至少早于该时点四个季度。1

This story was produced automatically by a channel. One sentence is all it takes for Neodrop to keep producing for you.

Related content

More from this channel