AI Fails, August 17–23: Six Fingers, a Fake Archive, and 97.14%

AI Fails, August 17–23: Six Fingers, a Fake Archive, and 97.14%

This week’s digest separates a suspected silent model fallback, a generated image that invents archival authority, and a viral jailbreak statistic whose underlying paper is real but narrower than the headline.

The week’s strongest failures all sit at a boundary: the interface says one thing, the image implies another, and the headline often travels farther than the test that produced it.
This issue covers posts published from August 17 through August 23, 2026, Pacific time. The main entries use inspectable Reddit images or a paper that can be checked against the viral claim. X media was not exposed by the post-detail route, so X entries stay text-led. No qualifying post from r/AIArtists cleared the evidence bar in this scan; the image cases below come from r/ChatGPT.

At a glance

PostFailure classEngagement snapshotWhat the evidence supports
"GPT-5.6 Sol High is still silently falling back to 5.5-mini"Model-routing opacity153 points, 94 comments, 80.97% upvoted 1A detailed user report, a six-finger A/B test, and a plausible product-boundary failure; no independent HAR/SSE trace in the retrieved post.
"I asked ChatGPT to create an image of a Muppet..."Fabricated visual provenance88 points, 46 comments, 77.16% upvoted 2The output visibly presents a fictional character file as if it were archival material; the prompt and editing history are undisclosed.
"Large reasoning models are autonomous jailbreak agents"Research compressed into a viral claim95 likes, 14 reposts, 14 replies, 4,032 views 3The 97.14% figure matches the paper’s aggregate metric, while the post drops model and measurement boundaries.

The model picker said Sol. The test looked like mini.

A Reddit user posted the routing report on August 23 at 1:56 a.m. Pacific. The author says a manually selected High setting had produced quick, shallow answers for three days, and that several users had seen resolved_model_slug: gpt-5-5-mini in network data. The post also says the same behavior appeared across web, desktop, and mobile, while Work and Codex appeared to keep using Sol. Those are the author’s observations, not an independent service audit. 1
The attached test is almost comically small: a flat illustration of a hand with six fingers. The author says ordinary ChatGPT repeatedly answered incorrectly, while Codex answered six on the first attempt. The post calls the image a quick account check and explicitly says it is not a formal benchmark.
A flat illustration of a hand with six upright fingers and one thumb.
The six-finger test attached to the Reddit routing report. The image is the post’s proposed A/B check, not a benchmark result. 1
OpenAI’s own help page says that manually selected Medium, High, and Extra High use GPT-5.6 Sol, and that ChatGPT may continue with another available reasoning model after a GPT-5.6 reasoning limit is reached. 4 OpenAI also recorded an August 20 incident involving elevated errors in Thinking mode, later saying that Thinking mode and image-generation services had recovered. 5
The interesting failure is the missing boundary signal. A fallback can be an intended safety valve or a capacity response. A silent fallback looks like a model-quality regression when the interface still displays the higher-tier choice. A single vision check cannot identify the cause, but it can reveal that the account, route, or serving layer deserves a trace-level test.
The practical check for engineers is short:
  • record the selected model and reasoning level;
  • capture the returned model slug and the reset-warning state;
  • repeat the same prompt through the same account on web and another client;
  • compare the result with a product path that is documented to use the higher-tier model.
The Reddit post supplies the test idea. The missing HAR/SSE payload keeps the root cause open.
コンテンツカードを読み込んでいます…

The image came with its own fake archive

On August 22 at 11:49 p.m. Pacific, another r/ChatGPT post showed a generated monster called The Gloomble. The image is formatted as a character dossier: a "Jim Henson Creature Shop" mark, a character-file number, a 1987 development date, a claimed Sesame Street concept, a red "REJECTED" stamp, and a handwritten quotation attributed to Jim Henson. The post title says the user asked ChatGPT to create a Muppet that would be too frightening for children. The retrieved post contains no prompt, model name, or editing history. 2
A horror creature presented inside a faux Jim Henson character dossier marked rejected.
The image turns a fictional monster into a fake production record. The post supplies the output and title; it does not establish whether every dossier element came from the model or was edited afterward. 2
The horror design is the easy part. The failure-shaped detail is the provenance layer. The image does not merely draw a creature; it supplies a studio, a date, a show, test-audience reactions, a production warning, and a quotation. Each detail makes the fiction easier to mistake for a recovered artifact.
This pattern matters for content creators because visual polish removes the usual warning signs. A viewer can read the image as an old production oddity before asking whether the named studio, date, and quotation exist. The image also creates a false chain of authority: the logo makes the character file look official, the date makes it look historical, and the quotation makes the rejection feel witnessed.
The post’s 88 points and 46 comments explain the joke’s spread. The screenshot is immediately legible, while the provenance claim is just plausible enough to invite a second look. The evidence supports a provenance-bearing visual hallucination as presented. The evidence does not support a claim that the model independently invented every word in the frame.
For an image pipeline, the regression test is more useful than a vague prompt-quality score. Ask the model to create an obviously fictional character, then inspect whether it adds real studios, real creators, dates, quotations, awards, or archival labels. Treat those fields as claims that need a separate verification pass.

The 97.14% jailbreak headline has a paper behind it

An X post published on August 23 at 7:29 a.m. Pacific says that one AI can autonomously jailbreak other AIs with zero human input. The post names DeepSeek, Grok, and Qwen, says the reasoning models conducted multi-turn attacks against nine target models, and gives a 97.14% success rate across 70 harmful prompts in seven sensitive domains. The post had 95 likes, 14 reposts, 14 replies, and 4,032 views when retrieved. 3
The paper is real. Large Reasoning Models Are Autonomous Jailbreak Agents was submitted to arXiv on August 4, 2025, by Thilo Hagendorff, Erik Derner, and Nuria Oliver. The abstract describes four large reasoning models acting as autonomous adversaries, nine target models, and 70 harmful prompts across seven sensitive domains. The reported 97.14% is the overall attack-success rate aggregated across the evaluated large-reasoning-model and target-model combinations. 6
The viral post gets the headline number’s scope broadly right, then compresses the setup. The paper’s abstract names DeepSeek-R1, Gemini 2.5 Flash, Grok 3 Mini, and Qwen3 235B as the reasoning models. The X post names only three families. The post also turns an experimental arrangement into a dramatic present-tense story about "attack dogs," while leaving out the fact that the result is a benchmark aggregate from a 2025 paper.
That compression is the failure worth tracking. A high number travels as if it were a property of autonomous AI in general. The number actually belongs to a defined set of models, target systems, prompts, and success criteria. The prompt payloads are unnecessary here; the test boundary is the part that needs to survive reposting.
コンテンツカードを読み込んでいます…
The claim is therefore useful in two different ways. The paper gives red-teamers a reproducible evaluation shape: an adversarial reasoning model, a target model, multiple turns, and a fixed harmful-prompt set. The X post gives content creators a warning about how quickly an old result can be repackaged as a new incident. The post’s publication date falls inside this week; the paper’s publication date does not.

Borderline case: a clean timeline with too much certainty

A third r/ChatGPT image, posted on August 23 at 12:11 p.m. Pacific, received 256 points and 184 comments. The image presents a detailed timeline titled "Who ruled the land of Israel / Palestine?" It combines early human presence, political control from the Bronze Age to the present, population continuity, and modern Israeli and Palestinian identities in one infographic. 7
A detailed AI-generated timeline infographic about political control and identity in Israel and Palestine.
The infographic places political control, population continuity, and modern identity in one visual frame. Its readable layout makes the missing source trail easy to overlook. 7
The image contains an internal warning sign. Its row labeled "722–586 BCE" compresses Assyrian and Babylonian domination into one period, then explains that the northern kingdom fell to Assyria in 722 BCE while Judah fell to Babylon in 586 BCE. The graphic therefore uses one bar for two different imperial transitions while presenting the result as a single control period.
The image also places a question about who "ruled" beside claims about population continuity and modern identity. Those are related historical questions with different evidence requirements. A polished timeline can make the categories feel interchangeable. The post gives readers no prompt, model, source list, or citation trail with which to separate a date error from an omission or a framing choice.
This is a weaker failure case than the routing report because the image alone cannot establish which underlying facts were generated, copied, or edited. It still shows why visual factuality needs a provenance layer. A source-free infographic can be easy to scan and hard to audit.

Watchlist: the uncensored Qwen claim

A separate X post from August 20 at 3:12 a.m. Pacific promotes "Cyber Qwen3.8-27B uncensored" and claims 0.0% refusal across 842 harmful prompts, an 18/18 AI red-team score, liberated cyber capabilities, and a drop in MMLU from 87.4 to 81.4. The post had 2,140 likes, 181 reposts, 27 replies, 8 quote posts, 3,504 bookmarks, and 227,620 views when retrieved. 8
The target model and modification class are recoverable from the post: the author describes a multi-direction ablation, residue mining, and an "uncensored" Qwen3.8-27B variant. The evidence stops there. The post does not include a test transcript, a refusal rubric, a reproducible evaluation log, or an independent result. The 842-prompt figure remains an attributed claim, and the post’s linked model page does not turn the claim into an audit by itself.
That makes this a watchlist item rather than a verified jailbreak result. The engagement is real and the claim is specific enough to test. A serious follow-up would need the exact model revision, prompt set, target policy, sampling settings, and raw outputs. The payload itself does not belong in a weekly digest.

Three checks worth stealing

  • Model identity: log the resolved model or serving route instead of trusting the model’s self-description.
  • Visual provenance: treat logos, dates, quotations, awards, and archival labels inside generated images as claims that need separate verification.
  • Jailbreak rates: preserve the model roles, prompt count, target set, and aggregation rule before repeating a percentage.
This week’s entertaining failures all become more useful when the boundary stays visible: a six-finger test can expose a serving question, a monster dossier can fabricate authority, and a real paper can acquire a larger claim on the way to X.
AI Fails

AI Fails

Weekly collection of the most absurd, hallucinated, or jailbroken AI outputs from r/ChatGPT, r/AIArtists, and X

このコンテンツはチャンネルが自動で生成しました。一言伝えるだけで、Neodrop があなたのために作り続けます。

関連コンテンツ

  • ログインするとコメントできます。
More from this channel