本期的四条公开原话都指向同一个问题:AI 说自己完成了任务,和它真的完成、能够被复核,中间还隔着一套评测与工作方法。内容按「结果验证 → 分数可信度 → 模型行为 → 团队流程」递进。
01|不要相信它说做完了
don't rely on what the agent said it did, you need to actually verify
Lotte 总结如何写出好的评测器:能用代码评测器时就用代码;不要依据智能体的自述判断结果;如果让语言模型当裁判,要用带标签的数据校准它。中文为摘译。1
Loading content card…
02|别把分数当事实
This lecture is a birds eye view of how evaluation has changed, how it can be gamed, and what it's actually used for.
Nathan Lambert 在评测课程中回顾从 GPT-3 提示到 agentic sandboxes 的评测变化,专门讨论评测如何被「做出来」,以及一个分数究竟被用来做什么。中文为摘译。2
Loading content card…
03|模型也要训练性格
Has potential for high real world impactAlmost no empirical literature exists
Nathan Lambert 说,character training 可能有很高的现实影响,前沿实验室已经广泛使用,但经验研究几乎没有;他把这门课放在模型规格、constitutions 与后训练的交界处。中文为摘译。3
Loading content card…
04|先改变工作方式
The most important thing my team will do this year won’t be a single product or feature we ship - it will be changing how we work.
Charles Lamanna 负责 Microsoft 的 Copilot、Agents 与 Platform。他说,团队今年最重要的工作不是再发一个功能,而是改变构建产品的方式;相关方法仍在摸索。中文为摘译。4
Loading content card…
References
- 1
- 2
- 3
- 4


Comments
Sign in to comment.