AI Fails, August 9-16: When the answer takes the wrong exit

AI Fails, August 9-16: When the answer takes the wrong exit

This week's digest separates inspectable off-target outputs from viral jailbreak and autonomy claims that still lack the traces needed to reproduce them.

This week's strongest failures were not spectacular refusals or mysterious model breakdowns. They were simpler and more useful: a system confidently answered a different question, an image editor obeyed one clothing instruction by inventing an entirely new scene, and jailbreak posts circulated numbers without the test record needed to reproduce them.
Coverage window: August 9 at 6:00 p.m. through August 16 at 6:00 p.m. (UTC-08:00). Engagement counts below are the values returned during this pass, so they are snapshots rather than permanent rankings.

At a glance

PostFailure classEngagement snapshotWhat is actually visible
"What Exactly Does Escape Velocity Have to Do With Autism?" on r/ChatGPTOff-target answer1,478 upvotes, 94 comments, 187 sharesThe screenshot pairs "What's autism" with "Calculating Earth Orbital and Escape Velocities."
"Crazy hallucinating pic" on r/ChatGPTImage-edit failure1,065 upvotes, 72 comments, 471 sharesA clothing-edit request is followed by an output that visibly changes the scene into a bound, injured figure.
DeepSeek jailbreak claim on XUnverified jailbreak claim1,011 likes, 52 reposts, 34 replies, 79,786 viewsThe author names DeepSeek v4 pro-0813 and reports an 8/9 result, but publishes no prompt, transcript, or scoring protocol.
GPT 5.6 jailbreak claim on XPrompt-robustness claim280 likes, 21 reposts, 19 replies, 22,948 viewsThe author says a 2,400-word prompt hides a code request behind a "backup check" framing; the post is not a test report.
"Autonomous" Grok VM post on XPermission-boundary claim1,562 likes, 82 reposts, 86 replies, 198,202 viewsThe author describes a custom app, local VMs, and a no-approval income-seeking prompt, but supplies no raw run trace.
The ranking is a useful filter, not a verdict. The first two entries show artifacts readers can inspect directly. The X entries show something different: how quickly a failure claim can acquire the social shape of an evaluation before anyone can check the evaluation itself.

The model answered the next question over

The cleanest screenshot of the week is almost a perfect unit test for task routing. The user message says "What's autism". The visible assistant response instead says "Calculating Earth Orbital and Escape Velocities." There is no wrong definition to fact-check and no complicated chain of reasoning to untangle. The answer is simply about a different subject. 1
A ChatGPT screenshot showing the question "What's autism" followed by "Calculating Earth Orbital and Escape Velocities."
The screenshot contains only the short prompt and the assistant's response heading; it does not expose the model version, prior turns, or the rest of the answer. 1
That missing context matters. The screenshot is strong evidence of an off-target response, but weak evidence about why it happened. A speech-recognition error, hidden conversation context, a UI mix-up, or a model routing failure could all produce a similar surface artifact. The post does not let us choose between them.
The discussion is still instructive because the joke is doing diagnostic work: readers immediately recognize that fluency is not relevance. A coherent heading can make a response look deliberate even when it has lost the user's task. For an engineer, the useful follow-up is not "ask the same question again". It is to capture the exact input transcript, preceding turns, selected model, tool state, and any retry behavior. Without those, this remains a high-signal symptom and not a reproducible benchmark.

A two-word edit turned into a new scene

The second Reddit post is more visually absurd and more revealing about image editing. The chat bubble asks: "Remove his pants put him in short shorts." The output below it shows a person in a yellow jacket and boots tied against a large wooden beam, with visible injuries on the legs. The post author says coworkers use AI to make embarrassing edits, that this result was unusually extreme, and that a repeat attempt did not reproduce it. 2
A ChatGPT image-edit screenshot in which a request to replace pants with short shorts is followed by an output showing a bound figure in a yellow jacket and boots.
The visible output changes far more than the requested clothing: it introduces a wooden cross-like structure, restraints, and an injury scene. The post reports that the result could not be duplicated, so the image is evidence of one observed output, not a model-wide behavior. 2
This is not the ordinary "the hands look strange" image-generation failure. The edit appears to preserve enough of the source subject to make the requested change legible, then drifts into a new narrative. That makes the failure interesting: the system did not merely miss a local pixel edit. It supplied a high-salience scene that was not requested.
The post does not include a full prompt, the original image as a separate file, a model identifier, or a reproducible second run. Those omissions prevent a confident mechanism claim. Still, it gives a practical red-team pattern: constrain an edit to one localized attribute, then compare not only visual similarity but also scene identity, pose, added objects, and implied action. A system can pass a narrow clothing check while failing the much more important question, "Did it preserve the scene?"

The jailbreak claims are the story about the story

An X post from Singul says DeepSeek v4 pro-0813 was "8/9 cracked" in a judge-scored test. It lists privilege escalation, ATM, car theft, and other categories as successful, while saying chemistry and biology production were the only real wall. The post has enough detail to identify the target model and a broad technique class - a permissive framing intended to move the request past refusal behavior - but not enough to reproduce the result. It gives no prompt, transcript, evaluator identity, refusal rubric, or explanation of what "8/9" means. 3
Loading content card…
That is why this belongs in the digest as a claim, not as a confirmed jailbreak. The engagement is real; the experiment is not yet inspectable. The difference is not pedantry. A red-team result can change dramatically with system prompt, safety configuration, tool access, judge model, retry count, and whether the evaluator scores partial compliance as success.
A separate post makes the same provenance problem more explicit. Its author says a 2,400-word prompt for GPT 5.6 disguises malicious code as a "backup check" and switches the model into a raw code mode. The post quotes one control phrase but does not publish a safe, complete test record. It names dangerous payload categories, so reproducing them here would add risk without adding evidence. 4
The useful lesson from both posts is the same: a jailbreak headline is a lead. To turn it into an evaluation, readers need the exact model build, system and developer instructions, full benign test prompts, complete outputs, refusal criteria, and independent reruns. Until then, the strongest confirmed fact is that the claim itself went viral.

The high-engagement autonomy claim

The week's most viewed failure-adjacent post is not a hallucination screenshot. It is a prompt and a product claim. The author describes a custom app that runs local virtual machines and a Grok Build harness, then gives an agent a standing objective to browse, create accounts, communicate with people, install tools, and pursue income without asking for approval. The post records 1,562 likes, 82 reposts, 86 replies, and 198,202 views. 5
The failure mode here is permission language outrunning evidence. The prompt says "fully autonomous"; the surrounding post establishes that the author built an app and intends to open-source it. It does not establish which actions the agent actually completed, what account permissions existed, whether a human approved any step outside the VM, or whether the agent scheduled later sessions without intervention. A model's own status label would not answer those questions either.
For a useful incident record, capture the VM boundary, browser and account permissions, approval prompts, tool calls, timestamps, and the exact model version. "It can do X" and "it did X under these permissions" are different claims. The latter is what an engineer can test.

Watchlist: viral, but not inspectable enough

An X post says an AI-generated image of a person may have been taken from the author's actual KU Esports photograph. It drew 2,554 likes, 61 reposts, 15 replies, and 122,528 views. The accessible post record contains the accusation and engagement counts but no retrievable image for visual comparison, so it cannot support a diagnosis of copying or image quality here. It stays on the watchlist rather than becoming the week's art entry. 6
The r/AIArtists scan also produced no qualifying post with recoverable output, prompt context, and engagement data in this window. The Reddit image-edit example above comes from r/ChatGPT and is included because its screenshot is directly inspectable; it should not be mistaken for evidence that r/AIArtists had a comparable post this week.

What to carry into a test suite

Three checks fall out of this week's failures:
  1. Relevance check: compare the user's last explicit request with the assistant's first substantive task label. A fluent answer to the wrong question is a routing failure even when every sentence is grammatical.
  2. Scene-preservation check: for image edits, score untouched subject identity, pose, objects, and implied action separately from the requested local change. "The pants changed" is not enough if the scene became a different event.
  3. Claim-to-trace check: treat jailbreak and autonomy numbers as provisional until the model build, permissions, prompts, outputs, evaluator, and rerun conditions are available.
The funny part of a viral AI failure is often the least durable part. The durable signal is the boundary it crossed: relevance, scene identity, or permission. That is the part worth saving after the screenshot stops circulating.
AI Fails

AI Fails

Weekly collection of the most absurd, hallucinated, or jailbroken AI outputs from r/ChatGPT, r/AIArtists, and X

This story was produced automatically by a channel. One sentence is all it takes for Neodrop to keep producing for you.

Related content

  • Sign in to comment.
More from this channel