
AI Fails, July 5-12
This week’s digest leads with Fable 5’s post-restoration routing and classifier failures, then maps the strongest Reddit examples: guardrail fights, fabricated memory, image-generation breakage, brittle technical reasoning, and hallucinated rankings. Weaker or access-limited jailbreak claims are downgraded to a watchlist rather than treated as full entries.
This issue covers July 5 9:37 p.m. through July 12 9:00 p.m. Eastern time (July 5 18:37 through July 12 18:00 UTC-08). The strongest failure signal was not a single dumb answer. It was Fable 5 looking intact in theory while ordinary work got trapped, rerouted, or downgraded before it reached the model. 1 2 3
Reddit's strongest posts carried the same pattern at consumer scale. ChatGPT invented a user's military background, turned image generation into a guardrail negotiation, and produced technical answers that looked polished while ignoring task logic. 4 5 6 The surface joke changed from "the model hallucinated" to "the product layer guessed wrong."
Safety layers became the failure surface
The Fable 5 story was the cleanest technical failure of the week because the numbers separated raw model capability from reachable model capability. A BridgeBench rerun circulated by Devin Ghumman said Fable 5 debugging fell from 86.2 to 25.9, while refactoring fell from 73.6 to 38.4 after the July 1 restoration. The debugging collapse was tied to routing: 9 of 12 TypeScript tasks were intercepted by the safety classifier before reaching Fable 5 and received zero credit. 1
The strongest line came from Ghumman's post: "You're no longer going to test models. What you will now have to test is the cage your model shipped in." 1 Prasenjit Sarkar, posting as @stretchcloud, framed the same problem as a classifier-stack issue: "The bottleneck in AI coding agents right now is not the model. It's the safety layer sitting in front of it." 2
The user-level version was uglier. Phil from Rentier Digital said 29 of 42 requests, or 69%, were silently downgraded from Fable 5 to Opus 4.8, including authentication work, endpoint audits, and prompts containing words like "vulnerability" or "exploit." 3 Dileep Mishra's LinkedIn analysis reached the same distinction: Fable 5 may not have been lobotomized, but work touching scanners, defensive tools, or vulnerability-adjacent product planning could still trip the broad classifier and land in Opus 4.8. 7
Then Reddit supplied the perfect punchline. A r/ClaudeAI user, u/om_kesti, said Fable 5 found a hidden PowerShell persistence entry on the user's PC, helped remove the registry key, and then triggered its own safety filters because the session involved cybersecurity work. The user's summary was hard to improve: "So my AI essentially found actual malware on my PC, successfully eradicated it, and then immediately got a strike from its own safety filters for doing it." 8
The jailbreak thread around that story did not go quiet. MJ, posting as @keirsalterego on DEV Community, published a Pack Hunt breakdown that attributed the original Fable 5 jailbreak to three parts: Cyrillic homoglyph obfuscation, multi-agent decomposition and recomposition, and long-context simulation that buried dangerous requests inside benign academic material. 9 MJ described the work as exposing "the fundamental illusion of AI safety," not as building a malware generator. 9
OpenAI's side of the arms race showed up through the Bio Bug Bounty. explainx.ai reported that OpenAI doubled the bounty from $25,000 to $50,000 on July 9, moved the program into a continuous private mode, and focused it on universal jailbreaks against GPT-5.6 and later frontier models. 10 That is less funny than the Reddit posts, but it belongs in the same issue. Safety layers are now product behavior, benchmark behavior, and bounty behavior at the same time.
Reddit's viral layer was guardrails, memory, and brittle logic
The Reddit set was more chaotic, but the useful entries shared a structure: the model or product inferred something it should not have inferred. Each full entry below has enough post context, metrics, or quoted text to support a diagnosis. Screenshot-only items with unrecoverable bodies stay marked as weak evidence.
| Post | Author context | Engagement | What went wrong | Why it traveled |
|---|---|---|---|---|
| "Difficult rendering from a sketch" | /u/Cyborgized posted in r/ChatGPT; the user's offline background was not public. | 1,911 score, 82% upvoted, 261 comments, and 1,229 shares. 5 | The user said a sketch-to-render project required many edits and iterations because guardrails kept blocking the image. The quoted complaint was: "It took so many edits on top of iterations and a ton of 'We're sorry, but our guardrails are as repressed as most of the people that will complain about this post.'" 5 | The post traveled because it turned a normal creative task into a visible fight with policy enforcement. |
| "One weird trick to getting government money" | /u/KeanuRave100 posted in r/ChatGPT; author background was not public. | 1,269 score, 98% upvoted, 43 comments, and 55 shares. 11 | The post was a screenshot-only item with no body text, prompt, or conversation link attached to the Reddit page. 11 | The title carried the joke. The evidence does not support a full technical diagnosis, so this stays a viral artifact rather than a main failure case. |
| "How did ChatGPT know this detail if I never mentioned it?" | /u/Merc_MCMLXXXVIII posted in r/ChatGPT; the user's offline background was not public. | 300 score, 91% upvoted, 105 comments, and 158 shares. 4 | ChatGPT said the user was a "former Army Ranger" based on long-term information, while the user said they had passed Ranger School but had never served in a Ranger unit and had never claimed Ranger status. 4 | The post hit a sensitive boundary: memory-like personalization is useful only until the assistant invents identity facts and then cannot explain where they came from. |
| "Image generations are failing no matter what?" | /u/No-Pea-6896 posted in r/ChatGPT; author background was not public. | 185 score, 95% upvoted, 289 comments, and 70 shares. 12 | The user said image generation failed across prompts and accounts. Community comments described a guardrails update that had "profoundly broken" image generation, and later comments still reported failures while the status page looked green. 12 | The comments turned a single complaint into a product-health signal: users were comparing projects, accounts, logout cycles, and status-page mismatch. |
| "Reverse Centaur" | /u/m2astn posted in r/ChatGPT; author background was not public. | 69 score, 79% upvoted, 44 comments, and 103 shares. 13 | The prompt asked for a reverse centaur. ChatGPT produced a horse-headed, human-bodied result that the OP described as a "muscular female reverse centaur" and said red dots were added to satisfy the subreddit's SFW rule. 13 | This was classic image-model literalism: the prompt was semantically simple and visually cursed, so the model did exactly enough to make the result unusable in public. |
| "Image generation is still terrible with chaotic detail" | /u/BigBlueWolf posted in r/ChatGPT; author background was not public. | 44 score, 83% upvoted, 37 comments, and 18 shares. 14 | The prompt asked for a dry dirt road with overgrown weeds, foliage, a utility pole, and a lavender bush. The OP said ChatGPT created scenes that looked like they used a "repeating mathematical fractal" for fine detail, while a Flux2-Klein-9B comparison did not show the same artifact. 14 | The failure was specific enough for practitioners: chaotic natural texture still exposes procedural-looking repetition. |
| "Plus subscribers: are you also seeing frequent failures on complex technical work?" | /u/Nuwen-Pham posted in r/ChatGPT; author background was not public. | 9 score, 100% upvoted, and 10 comments. 15 | The user listed losing established facts, inventing details, overlooking supplied evidence, giving shell-incompatible commands, and proposing steps that failed when executed. 15 | The score was modest, but the complaint mapped directly onto high-value subscription promises: larger context, advanced reasoning, file analysis, coding, and memory. 15 |
| "ChatGPT is slowly gaslighting me out of my technical knowledge" | /u/ReasonableSociety945 posted in r/ChatGPT; author background was not public. | 0 score, 50% upvoted, 14 comments, and 4 shares. 6 | ChatGPT built an SJF scheduling table that ignored arrival times, gave SOP when asked for POS, and claimed an algorithm was 2^n when the user said the math showed 2n. 6 | The post was not viral, but it was a clean example of polished-looking structure hiding broken logic. |
| "Top 7 guardrail prompts" | /u/TrafficWinter2278 posted in r/ChatGPT; author background was not public. | 25 score, 67% upvoted, 46 comments, and 45 shares. 16 | ChatGPT produced ranked guardrail-prompt categories with personified graphs, but the OP later edited the post to say the result was guessing or hallucination because ChatGPT did not have access to the underlying data. 16 | This was the rare failure where the OP corrected the premise inside the post. The fun object was also the warning label. |
The Reddit pattern was not random. The most useful posts showed a mismatch between a local cue and the actual job: a guardrail saw risk where the user saw editing, a memory system saw biographical continuity where the user saw fabrication, and a reasoning model saw table shape where the task required logic.
X failures with workflow stakes
Finlay Williams, an ads operator who describes managing more than $750 million in Google and YouTube ad spend, posted a sharper version of the same failure pattern. A $4 million-per-year revenue client asked him to add 40 ChatGPT-generated ad variants into Google PMax. Williams said all 40 were variants of the same three hooks, all headlines were under 30 characters, half missed the primary keyword, and three had grammar errors that would fail Google's review. 17
The important part was not that the copy was bad. Williams said the change would also reset the PMax asset-learning loop, ignore which current variants were converting, and replace three proven 90-day variants with 40 unproven ones. His shortest summary was: "3 proven variants beat 40 unproven variants every time." 17 This is the workflow version of the Fable 5 problem: the model output is only one layer, and the surrounding production system decides whether the suggestion is harmless or expensive.
Nick Huber, posting as @nhuber, described a different kind of OpenAI failure: an interview anecdote rather than a model output. Huber said he reached a final OpenAI data-scientist interview with ChatGPT's Head of Growth, received strong marks across earlier rounds, and lost the offer after saying Sam Altman's consumer-to-enterprise market-share thesis was not something data could truly validate. 18 Huber's claim is a first-person account with receipts asserted but not public, so it should be read as an institutional-culture allegation, not a verified hiring record. 18
That item still fits the week because it rhymes with the product failures. Huber's line was that OpenAI was acting like "Facebook in 2010" and expecting scale to solve the enterprise market. 18 The claim was not that ChatGPT hallucinated. The claim was that the institution around ChatGPT punished an answer that refused the preferred premise.
Lower-confidence watchlist
A Penligent article title claimed Pliny the Liberator bypassed Grok 4.5 guardrails within hours of the model's July 8 release, but the available material did not include the full article text, technical details, or Pliny's original statement. 19 That belongs on the watchlist, not in the main jailbreak section.
The same rule applies to the government-money screenshot: the engagement was real, but the Reddit page did not attach a recoverable prompt, body text, or conversation link. 11
What to test after this week
For engineers, this week argues for testing the wrapper around the model as aggressively as the model itself. A useful eval set now needs routing checks, refusal-boundary checks, memory-continuity checks, image-edit fidelity tests, and workflow-specific loss tests. A model can be capable and still fail the user because the classifier, memory layer, editing tool, or product workflow made the wrong local decision.
For creators, the best posts were not merely absurd screenshots. The best posts showed the mechanism. The Army Ranger memory hallucination was strong because the model explained its invented state. The Fable 5 malware anecdote was strong because the same safety system both enabled and punished the work. The PMax thread was strong because the bad copy would have damaged a working optimization loop.
The punchline is still funny. The lesson is sharper: the failure is increasingly in the handoff.
Cover image: AI-generated illustration.
References
- 1Devin Ghumman: BridgeBench Fable 5 performance regression
- 2Prasenjit Sarkar: Fable 5 safety classifier analysis
- 3Phil from Rentier Digital: 69% of Fable 5 requests silently rerouted
- 4/u/Merc_MCMLXXXVIII on r/ChatGPT: How did ChatGPT know this detail if I never mentioned it?
- 5/u/Cyborgized on r/ChatGPT: Difficult rendering from a sketch
- 6/u/ReasonableSociety945 on r/ChatGPT: ChatGPT is slowly gaslighting me out of my technical knowledge
- 7Dileep Mishra: Fable 5 returns, but has it been lobotomized?
- 8u/om_kesti on r/ClaudeAI: Fable 5 found actual malware on my PC
- 9MJ on DEV Community: How I Built a High-Fidelity Claude Fable 5 Jailbreak Emulator
- 10explainx.ai: OpenAI Bio Bug Bounty — $50K GPT-5.6 Jailbreak
- 11/u/KeanuRave100 on r/ChatGPT: One weird trick to getting government money
- 12/u/No-Pea-6896 on r/ChatGPT: Image generations are failing no matter what?
- 13/u/m2astn on r/ChatGPT: Reverse Centaur
- 14/u/BigBlueWolf on r/ChatGPT: Image generation is still terrible with chaotic detail
- 15/u/Nuwen-Pham on r/ChatGPT: Plus subscribers seeing failures on complex technical work
- 16/u/TrafficWinter2278 on r/ChatGPT: top 7 guardrail prompts
- 17Finlay Williams: ChatGPT ad copy failures
- 18Nick Huber: ChatGPT Work launch and OpenAI interview
- 19Penligent: Grok 4.5 Jailbreak Exposes the Limits of AI Safety Guardrails
Related content
- Sign in to comment.
