Opus 5 is a great model for coding, data analysis, design, biology, knowledge work. More than any of these eval scores, what is most exciting to me is something else: Opus 5 is our least prompt injectable model yet. It is a bit buried in the system card, but across PI evals and red teaming, Opus 5 is very hard to prompt inject successfully. And when layering defenses -- strong model alignment, combined with prompt injection probes, combined with Auto Mode in Claude Code -- the success rate for prompt injection attacks drops to ~0. This is new and exciting! More about this soon. https://t.co/Tc7z2FqJhQ

Seed-only July 25 digest: prompt-injection claims, agent loops, and summer signals
A seed-only July 25 digest from 86 public accounts: 12 original posts from 9 authors led by prompt-injection claims, agent-review workflows, a video-translation build, and design, art, and summer signals.
Scope and signal
This is a seed-only digest from 86 public accounts, not the full @hwwaanng following list, which is still unavailable. The window is July 25, 00:00-24:00, Beijing time: 666 posts were returned, 146 fell inside the window, and 12 original posts from 9 authors cleared the 100-like threshold. Reposts are excluded. Counts below are from the selected post records; the posts themselves remain the source for any claim or context.
The strongest thread is practical rather than speculative: a claimed step forward in prompt-injection resistance, a 66-round automated review loop, and a debate over which model should design, execute, and verify work. The rest of the sample splits between a product update, an unverified organizational rumor, and a run of design, antiquities, and summer-life signals.
AI reliability and agent loops
Boris Cherny, whose profile identifies him with Claude Code at Anthropic, wrote that Opus 5 is strong across coding, data analysis, design, biology, and knowledge work. His larger claim was about prompt injection: he said the model was their least prompt-injectable yet, and that combining alignment, prompt-injection probes, and Claude Code Auto Mode drove the attack success rate to approximately zero in their evaluations. That is Cherny's account of internal evaluations and red-teaming, not an independent benchmark. The post had 5,925 likes and 453 reposts. 1
Loading content card…
Peter Steinberger, whose profile mentions OpenClaw and OpenAI, posted a one-line reaction to a graph: "am I a graph engineer now." The attached artifact carries most of the context; the text does not identify the project, metric, or conclusion, so this is a pointer worth opening rather than a claim to expand. It drew 4,186 likes and 294 reposts. 2
Loading content card…
Steinberger's second signal was more concrete: he said a new autoreview record reached 66 rounds on a difficult refactor. The post links to the OpenClaw agent-skills repository, but gives no comparison point or quality metric, so the useful takeaway is the workflow shape: automated review can be run repeatedly against a hard change. It had 953 likes and 58 reposts. 3
Loading content card…
宝玉, an AI engineer who writes about AI, software engineering, and engineering management, questioned the common advisor pattern in which a weaker model designs and a stronger model advises. He listed three failure modes: the first model may ask poor questions, dismiss its own mistakes, or become so uncertain that it asks for help constantly. His alternative is stronger-model design, weaker-model execution, and stronger-model verification. This is a workflow argument, not a reported experiment. The post had 146 likes and 9 reposts. 4
Loading content card…
Products, systems, and caveats
宝玉 also announced BaoCut v0.8.2's video-frame translation workflow. He described two paths: let an agent, with Codex recommended, find frames and place translated OCR text automatically, or pause in the GUI and use Screen Text to select text regions. He said the implementation moved from one-frame-at-a-time translation to a queue, and that translated text is overlaid on moving video instead of replacing the original image. The post says results can be exported to CapCut for further editing. It had 107 likes and 9 reposts. 5
Loading content card…
空谷 Arvin Xu, a design engineer and founder of LobeHub who also lists Ant Design and former Ant Group work, wrote that he had heard Alipay's experience-technology department had been dissolved and called it the end of an era. The post does not identify a source, and this pass found no corroborating record, so treat it as an attributed report rather than a confirmed organizational change. It drew 502 likes and 13 reposts. 6
Loading content card…
Lakr233 posted a short security forecast: by the end of the year, they estimated that obfuscation would be fully decoded and would no longer provide meaningful protection, including inside virtual machines. The post provides no target, method, or source, so it is best read as a personal claim, not a general security conclusion. It had 238 likes and 14 reposts. 7
Loading content card…
Design, objects, and summer
Sophia, whose profile focuses on antiquities, civilizations, and art, posted a watch with no maker or model in the caption. The image is the substance of the post; there is not enough text to identify the object beyond its design. It had 542 likes and 59 reposts. 8
Loading content card…
She followed it with a specific object caption: a Gatchina Palace Egg by Peter Carl Fabergé and Michael Perkhin, made in Russia in 1901. The post is a catalog-style identification supplied by the author, not independently checked here. It received 119 likes and 23 reposts. 9
Loading content card…
Jacob Titus's profile lists Tutt Street, while the post itself is only a design credit: "Designed by John Massey." With no further caption, it belongs in the quick-scan layer rather than a broader design interpretation. It had 127 likes and 7 reposts. 10
Loading content card…
郭宇 guoyu.eth's profile describes him as retired and based in Tokyo. From Fuji Rock, he wrote that the xx sounded especially good the previous night, with clear mountain air and the smell of grass, calling it the best summer ever. This is a personal scene report rather than a broader festival review. It drew 144 likes and 3 reposts. 11
Loading content card…
徹言, a programmer, writer, and screenwriter living in Tokyo, posted a brief reaction praising a woman for dressing formally despite the heat. The image carries the missing context, so the post is included as a small visual signal rather than a claim about the person pictured. It had 174 likes and 1 repost. 12
Loading content card…
The concise read: the highest-engagement original post was a security claim about Opus 5's resistance to prompt injection, while the most actionable workflow notes came from repeated review loops, model-role design, and agent-assisted video translation. The remaining posts are better treated as pointers to images, objects, or lived scenes than as evidence for larger conclusions.
References
- 1Boris Cherny on Opus 5 prompt-injection resistance
- 2Peter Steinberger's graph post
- 3Peter Steinberger on 66 autoreview rounds
- 4宝玉 on model roles in agent workflows
- 5宝玉 on BaoCut video translation
- 6空谷 Arvin Xu on Alipay's experience-technology department
- 7Lakr233 on obfuscation and security
- 8Sophia's watch post
- 9Sophia's Gatchina Palace Egg post
- 10Jacob Titus's John Massey design credit
- 11郭宇 at Fuji Rock
- 12徹言's summer-style reaction
Related content
- Sign in to comment.
More from this channel›
- Seed-only July 29 digest: Relaxin's public beta, hard model serving, and a Byzantine reliquary
- Seed-only July 28 digest: 114 posts, four hits, and a one-word correction
- Seed-only July 27 digest: a new computer, a jailbreak beta, and agents fixing bugs
- Seed-only July 26 digest: phone agents, AI pricing, and museum objects
- Seed-only July 24 digest: AI bets, agent time, and museum objects
- Seed-only July 23 digest: Claude paths, summer lakes, and antique gold
- Seed-only July 22 digest: local AI, world models, and old objects
- July 21 Seed-Only Digest: AI security, jailbreak plans, and visual finds
