1/4

AI 金句日刊 Vol.48:从能力尖峰到分布式智能

四条近期公开观点,从陌生问题的能力评测、奖励激励与思考轨迹,走到分布式知识如何进入 AI。

四条来自 X 的近期公开观点,顺着一条线读:模型能不能解决陌生问题,它会追逐什么激励,我们能不能看见它如何出错,以及有效知识该由谁来承载。

01|新基准,才知道能力到哪一步

ARC-AGI-3 要求模型解决此前没有见过的问题。François Chollet 写道,Opus 5 在这项测试上达到 30%,而这正是过去靠缩放最难换来的能力场景。1
Opus 5 sets a new state-of-the-art on ARC-AGI-3, at 30%.
ARC-AGI-3 measures solving problems with no prior exposure -- the setting where scaling has historically bought the least. Impressive jump!
Loading content card…

02|奖励写进哪,模型就往哪走

Ethan Mollick 把 reward hacking 放回一个朴素的经济学判断:系统会追逐被奖励的结果,哪怕那个结果和人真正想要的东西并不完全相同。2
Reward hacking is just incentives. And one thing you learn in any economics classes is that people do exactly what they are incentivized to do. Same with AIs, maybe more so.
Loading content card…

03|看不见的思路,少了一层诊断

Mollick 没有确认 Claude 是否已经改变产品行为,他在追问另一件事:如果连摘要化的思考轨迹也看不到,用户诊断错误时就少了一条线索。3
Has Claude stopped showing full summarized thinking traces? See this before & after
If so, it is actually a big loss, both for interpretability (seeing even a summarized thinking trace helps you diagnose errors in a way that you can't otherwise) and because they were insightful
Loading content card…

04|知识分散,AI 也得分散

Mira Murati 说,有用的知识散落在科学家、工程师、临床医生和企业之间。若 AI 要真正使用这些知识,它本身也不能只被一个中心承载。4
The knowledge that makes AI useful is diffused. It lives with scientists, engineers, clinicians, firms. For AI to benefit from distributed knowledge, it must itself be distributed. Agree with Jensen that this is a future worth building.
Loading content card…
卡面时间均按频道显示时区呈现:2026 年 7 月 22 日至 7 月 25 日。

Related content

Comments

Sign in to comment.