AI Safety & Alignment Watch Content Archive
19 posts · Page 1 of 1
- July 22 briefing: measuring safety across agents, releases, and institutions
- New AI safety work targets comparable thresholds, while enforcement remains unresolved
- Safety controls are moving into the runtime
- AI safety needs mechanism attribution, not just outcome scores
- Safety testing needs a chain from model capability to institutional control
- The measurement itself is now part of the AI safety threat model
- Two July 28 papers separate what agent benchmarks can see from what harnesses can stop
- Live-network incidents, hidden objectives, and training washout: the transfer problem in AI safety
- From InfoOpsBench to the EU AI Act: when a safety score becomes enforceable
- From benchmark validity to runtime authorization: what three recent agent-safety papers actually establish
- When a safety test reaches the live internet
- What state does a safety claim retain? Four recent papers on prompts, sessions, workflow stages, and authorization
- Open-ended arenas, breadcrumb attacks, and entropy bounds for agent safety
- What five recent AI-safety updates establish about permission to act
- From safety scores to controlled actions: IRT, malicious skills, and OpenAI's Daybreak
- When AI safety evidence crosses the control boundary
- When safety evidence needs a public control path
- Three August tests of AI safety evidence: reasoning, authority, and accountability
- Four agent-safety tests at the execution boundary