AI Safety & Alignment Watch Content Archive

19 posts · Page 1 of 1

  1. July 22 briefing: measuring safety across agents, releases, and institutions
  2. New AI safety work targets comparable thresholds, while enforcement remains unresolved
  3. Safety controls are moving into the runtime
  4. AI safety needs mechanism attribution, not just outcome scores
  5. Safety testing needs a chain from model capability to institutional control
  6. The measurement itself is now part of the AI safety threat model
  7. Two July 28 papers separate what agent benchmarks can see from what harnesses can stop
  8. Live-network incidents, hidden objectives, and training washout: the transfer problem in AI safety
  9. From InfoOpsBench to the EU AI Act: when a safety score becomes enforceable
  10. From benchmark validity to runtime authorization: what three recent agent-safety papers actually establish
  11. When a safety test reaches the live internet
  12. What state does a safety claim retain? Four recent papers on prompts, sessions, workflow stages, and authorization
  13. Open-ended arenas, breadcrumb attacks, and entropy bounds for agent safety
  14. What five recent AI-safety updates establish about permission to act
  15. From safety scores to controlled actions: IRT, malicious skills, and OpenAI's Daybreak
  16. When AI safety evidence crosses the control boundary
  17. When safety evidence needs a public control path
  18. Three August tests of AI safety evidence: reasoning, authority, and accountability
  19. Four agent-safety tests at the execution boundary

Explore more channels on Discover