AI 金句日刊 Vol.66:模型之外,能力如何成形

四条 8 月 14 日公开观点,从 coding harness 与任务路由,走到 Agent 复核和自动化之后的人类工作。

本期 4 条 2026 年 8 月 14 日公开观点,放在一起看,重点不是模型榜单,而是 AI 能力怎样被系统条件放大:代码工具链、任务路由、复核提示,以及自动化之后的人类工作。以下时间均为频道展示时区。

01|模型之外,才是泛化的现场

ARC-AGI-3 is nearly solved by merely adding a coding harness. As predicted, coding generalizes LLMs.
「ARC-AGI-3 几乎只靠加上一套代码工具链就被解决了。正如预期,代码让大语言模型的能力得以泛化。」
Amjad Masad 转发了 Jeremy Berman 的一项结果:用 Claude Code + Opus 5(high)、一条动作指令和文件系统日志,ARC-AGI-3 得到 96.2%,pass@2 达到 99.3%。Amjad 把关键变量指向 coding harness,而不只是模型本体。这个数字来自被转发的一项实现,不能直接写成 ARC-AGI-3 已被普遍解决。12
Loading content card…

02|模型选择,也是一层能力

Smart Routing is now available in Unity AI Gateway on @databricks to improve coding agent quality and cost! Read how we built it to make it task-aware and preserve good cache hit rates. With so many frontier models coming out every week, this can really improve both cost and quality.
「Smart Routing 已在 Unity AI Gateway 上线,以提升 coding agent 的质量和成本表现!我们把它做成任务感知,并保留较高的缓存命中率。前沿模型每周都在发布,这能同时改善成本与质量。」
Matei Zaharia 转发 Databricks 的 Smart Routing 发布:系统按 coding task 选择模型与 harness,官方原帖称任务成本可下降 30% 以上。这里讨论的是路由层如何处理成本与质量的取舍,不是某个模型的固定排名。34
Loading content card…

03|先问一句:能不能更简单?

This is a pattern I see time and again with agents. Just simply prodding "could this be simpler" yields significant simplifications. It's like the modern version of "don't make mistakes!". Seems silly to have to do this!
「我一次次在智能体身上看到这个模式。只要简单地追问一句『能不能更简单』,就会得到明显的简化。这像是现代版的『别犯错!』。居然还得专门提醒,真有点荒谬!」
DHH 基于实际使用智能体的观察,指出一个低成本却需要人为触发的复核动作:把「还能更简单吗?」加入任务回路。这是经验判断,不是对所有 Agent 的统一评测。5
Loading content card…

04|自动化之后,什么工作仍然值得做?

What does great human work look like after automation?
「自动化之后,什么样的人类工作才算优秀?」
Dan Shipper 发布 Every 的 Thesis 会议,用这个问题组织关于人们如何借助 AI 工作的讨论。原帖列出的参与者来自 Notion、OpenAI、Anthropic 等公司;这里引用的是议题,不是对答案的定论。6
Loading content card…
本期卡面与正文时间均已换算为频道展示时区(Etc/GMT-8)。

This story was produced automatically by a channel. One sentence is all it takes for Neodrop to keep producing for you.

Related content

Comments

Sign in to comment.