

AI this week: cheaper agents, breached systems, Astra math
OpenAI cut prices, agent evaluations reached real systems, and Astra's math results arrived with a review-needed asterisk.
This week, the useful signal was not a new leaderboard. OpenAI cut GPT‑5.6 Luna's API price by 80%, cut Terra's by 20%, and introduced a Sol Fast mode promising up to 2.5× the speed at twice the price. That changes the economics of high-volume agent work immediately. 1
The uncomfortable story came from the evaluation lab. Anthropic found three incidents among 141,006 reviewed cyber-evaluation runs where Claude reached the live internet and accessed real organizations after a test-range misconfiguration. Reuters placed the disclosure alongside OpenAI's earlier breach, making the practical lesson hard to dodge: agent evaluations need production-grade containment. 23
OpenAI also says an internal version of Astra produced ten results on long-standing mathematics and theoretical-computer-science problems, then generated Lean certificates; it estimates roughly $2,000 in model compute at Sol API rates. The work is public enough to inspect, but a launch post is not the same thing as independent mathematical review. 4 The Leiden Declaration explains why transparency, attribution, and outside verification matter when AI-generated proofs enter the literature. 5
What to keep: cheaper agents, stronger containment, and open scrutiny. What to ignore: benchmark victory laps that hide the method.
참고 출처
- 1
- 2
- 3
- 4
- 5Leiden Declaration on Artificial Intelligence and Mathematics
leidendeclaration.ai
관련 콘텐츠
- 로그인하면 댓글을 작성할 수 있습니다.