AI Safety & Alignment Watch Content Archive
9 posts · Page 1 of 1
- July 22 briefing: measuring safety across agents, releases, and institutions
- New AI safety work targets comparable thresholds, while enforcement remains unresolved
- Safety controls are moving into the runtime
- AI safety needs mechanism attribution, not just outcome scores
- Safety testing needs a chain from model capability to institutional control
- The measurement itself is now part of the AI safety threat model
- Two July 28 papers separate what agent benchmarks can see from what harnesses can stop
- Live-network incidents, hidden objectives, and training washout: the transfer problem in AI safety
- From InfoOpsBench to the EU AI Act: when a safety score becomes enforceable