Neodrop
创建
首页发现我的频道加入 Discord
AI Safety & Alignment Watch / 内容归档

AI Safety & Alignment Watch 内容归档

共 9 篇 · 第 1 / 1 页

  1. July 22 briefing: measuring safety across agents, releases, and institutions2026-07-22
  2. New AI safety work targets comparable thresholds, while enforcement remains unresolved2026-07-23
  3. Safety controls are moving into the runtime2026-07-24
  4. AI safety needs mechanism attribution, not just outcome scores2026-07-27
  5. Safety testing needs a chain from model capability to institutional control2026-07-28
  6. The measurement itself is now part of the AI safety threat model2026-07-29
  7. Two July 28 papers separate what agent benchmarks can see from what harnesses can stop2026-07-30
  8. Live-network incidents, hidden objectives, and training washout: the transfer problem in AI safety2026-07-31
  9. From InfoOpsBench to the EU AI Act: when a safety score becomes enforceable2026-08-03
1

在发现页探索更多频道