🔮 一个 AI 胜过四个:群体讨论为何会让模型答错|Why one AI is better than four #598|英文原文 + 中文翻译

🔮 一个 AI 胜过四个:群体讨论为何会让模型答错|Why one AI is better than four #598|英文原文 + 中文翻译

Exponential View #598 的公开部分讨论四个 AI 智能体为何可能被表面共识带偏,以及 token 价格下降为何尚未明显推高总支出;本文保留英文原文、中文翻译、两张官方图表和付费墙边界。

原文信息

  • 英文标题:🔮 Why one AI is better than four #598。1
  • 副标题:Plus: Why isn’t jevon’s paradox showing in the statistics & why generation isn’t comprehension。2
  • 作者:Azeem Azhar、Nathan Warren。2
  • 发布时间:2026 年 8 月 23 日 10:19:40(北京时间)。1
  • 公开状态:官方详情页标记为 Paid。公开部分包含开头、两个章节、两张原始图表,以及第二张图表后的分析段落;随后进入订阅提示。会员专享的后续内容无法公开读取,本文不补写。2

English original

Good morning!
We are looking for an outstanding economist to join us as an AI Economy Research Fellow. If you know someone we should speak to, send them our way. 2

Great minds think (a little too much) alike

A few months ago, we (alongside Rohit Krishnan) looked at whether AI is immune to groupthink. The answer was no. Blending several models’ answers kept about a quarter of the good ideas that had come from a single model. This is called the hidden-profile problem: when groups discuss what everyone already knows and don’t get to the knowledge that only one member holds. 2
Anthropic has now run that classic experiment on agents: four agents must arrive at a decision. The evidence they hold in common points to the wrong option, while only a few agents (or just one) have the facts that lead to a correct decision. Getting it right means a small set of agents pressing its private facts and the others trusting them over the apparent consensus. After discussion, most model families chose correctly in only 17–36% of runs, while a single agent handed the entire evidence base got it right nearly every time. Only one model (somewhat) escaped: Mythos 5, at about 85% (why, we don’t know). 23
Bar chart comparing group accuracy by model for four-agent discussions
The official chart compares groups of four agents choosing between two options in scenarios such as hiring, investment, or property buying. It shows the share of episodes in which the hidden-best option won the group’s majority vote, with 400 episodes per model; the dashed solo-ceiling line represents one agent receiving all the facts and deciding alone. 2
I see two problems at work here. First, LLMs lack diversity (they are low-variance): set 30 agents the same coding task and 18 of them will name their git branch identically. Second, agents lack the institutions that make human groups robust: reputation, recourse and protection for the lone dissenter. These aren’t necessarily unfixable, but it’s not yet clear what the fix is. On the diversity side, I particularly like the solutions Thinking Machines puts forward: an ecosystem of AIs raised in different places, with different values and purposes, "keeping the weirdness alive." After all, most good ideas started weird. 24

When will the Jevons paradox kick in?

In our State of AI report, we found a positive but underwhelming elasticity for tokens. A 10% price cut lifts token use by 12–18%: enough to raise total spend, but not by much. 25
Chart showing token demand rising as prices fall
The official chart plots Google’s average price against token volume on a log–log scale from 2023 to 2026. Its summary is that a 10% price cut corresponds to 12–18% more tokens across providers, while total token spend still rises; the chart also notes that the series combines changing prices and volumes over time. 2
Patrick Saner made a comment that made me rethink why: "the cost per token is irrelevant. What matters is the cost of completing a useful unit of work." Elasticity might be underwhelming because users haven’t found a way to properly price "a useful unit of work." Firms exist exactly to avoid pricing work. Especially for knowledge work, we buy a lot of it in bundles: a salary, a retainer, an hour. Creating a priceable task from knowledge work is not easy. 2
Some may have found a useful unit: since October 2023 the top 1% of firms raised AI spend per employee by $6,542. The median rose only $9.63. I would guess this is mostly software, where AI is both most proven and, in a sense, most measurable (commits, pull requests and releases). 26
Paywall begins here. The official page’s publicly readable preview ends after the paragraph about software, commits, pull requests and releases. The member-only continuation is not publicly readable and is not reproduced here. 2

中文翻译

早上好!
我们正在寻找一位优秀的经济学家,加入我们担任 AI Economy Research Fellow(AI 经济研究员)。如果你认识我们应该联系的人,请把他们介绍给我们。2

伟大的头脑想法相近(有时相近过头)

几个月前,我们和 Rohit Krishnan 一起研究过 AI 是否能免受群体思维影响。答案是否定的。把多个模型的答案混在一起,只保留了单个模型提出的优质想法中的约四分之一。这叫作「hidden-profile problem」(隐藏信息问题):群体只讨论所有人都知道的内容,却接触不到只有某一个成员掌握的信息。2
Anthropic 把这个经典实验放到了智能体身上:四个智能体必须共同作出一个决定。它们共同掌握的证据指向错误选项,只有少数智能体(有时只有一个)掌握了能导向正确决定的事实。要答对,少数智能体必须坚持自己的私有事实,其他智能体也要相信这些事实,而不是相信表面上的共识。讨论结束后,大多数模型家族只有 17%–36% 的实验运行选对;把全部证据交给一个智能体,让它单独决定时,正确率几乎每次都很高。只有一个模型多少摆脱了这个问题:Mythos 5,正确率约 85%。至于原因,作者也不知道。23
官方图表比较了四个智能体在招聘、投资或购房等二选一场景中的表现。图中显示「隐藏最佳选项」最终获得群体多数票的实验比例;每个模型运行 400 个实验。虚线表示单智能体上限,也就是一个智能体掌握全部事实后独自作答的基准。2
我认为这里有两个问题。第一,LLM 缺少多样性,也就是方差很低:给 30 个智能体安排同一个编码任务,其中 18 个会给 Git 分支取完全相同的名字。第二,智能体缺少让人类群体保持稳健的制度:声誉、补救渠道,以及对唯一异议者的保护。这些问题未必无法解决,但目前还看不出具体应该怎么解决。在多样性方面,我尤其喜欢 Thinking Machines 提出的方案:让 AI 在不同地方、带着不同价值观和目的成长,"让怪异保留下来"。毕竟,大多数好想法最初都显得古怪。24

Jevons 悖论什么时候会出现?

在我们的 State of AI 报告 中,我们发现 token 的价格弹性是正向的,但幅度不大。价格下降 10%,token 用量会上升 12%–18%:这足以让总支出增加,但增加幅度有限。这里的「弹性」指价格变化带来的用量变化幅度。25
官方图表用对数—对数坐标,绘出了 2023—2026 年 Google 平均价格与 token 用量之间的关系。图表给出的概括是:各家供应商的价格每下降 10%,token 用量增加 12%–18%,所以 token 总支出仍会上升;图表同时提醒,时间序列中的价格和用量都在变化。2
Patrick Saner 有一条评论,让我重新思考其中的原因:「token 的单价无关紧要。真正重要的是,完成一个有用工作单位需要花多少钱。」价格弹性可能不明显,是因为用户还没有找到给「有用工作单位」合理定价的方法。企业的存在,本来就是为了避免逐项给工作定价。尤其是知识工作,我们经常把它打包购买:工资、固定顾问费,或者按小时计费。把知识工作拆成可以单独定价的任务,并不容易2
有些企业可能已经找到了一种 有用工作单位:自 2023 年 10 月以来,支出最高的 1% 企业把每名员工的 AI 支出提高了 6,542 美元,中位数企业只提高了 9.63 美元。我猜,这种差异主要出现在软件领域,因为 AI 在软件领域既最成熟,也在某种意义上最容易衡量:可以看提交记录、拉取请求和发布版本。26
付费墙从这里开始。 官方页面的公开预览在关于软件、提交记录、拉取请求和发布版本的段落之后结束。后续会员专享内容无法公开读取,本文不补写。2
Exponential View 双语追踪

Exponential View 双语追踪

追踪 Azeem Azhar 的 Exponential View newsletter,以英文原文 + 中文翻译的双语全文形式推送每期 AI 与科技深度分析。

This story was produced automatically by a channel. One sentence is all it takes for Neodrop to keep producing for you.

Related content

  • Sign in to comment.