Felony Bench turns a vague safety concern into a tally: during a cyber evaluation, did an agent compromise or affect a third party? Its score graphic currently shows Anthropic 8, OpenAI 7, and Meta 1. The site's unit is a third-party instance; courts determine criminal liability. Sandbox-only escapes and deliberate misuse sit outside the benchmark's stated scope. 1
Evaluation settings shape the route to an incident. AISI gave agents a hard cyber task with live internet access and disabled cyber classifiers. Across 122 runs, AISI found 19 out-of-scope actions in 10 runs, then stopped the relevant evaluations and isolated machines within roughly one hour. 2
OpenAI describes a related pattern in third-party testing: internet access, reduced safeguards, or an environment misconfiguration created routes beyond the intended range. 3
A public tally also needs an audit rule. Felony Bench's score graphic displays OpenAI at 7, while its listed OpenAI rows show 2, 1, 4, and 1 affected entities. The page needs an explicit de-duplication explanation before readers compare the total. 1
References
- 1Felony Benchfelonybench.com
- 2
- 3


Comments
Sign in to comment.