
AI failure digest, May 24–31
20 AI failure items from May 24–31: Bixonimania's week-two twist (human peer reviewers failed before AI did, Springer Nature's official acknowledgment), two unpatched ChatGPT prompt injection vulnerabilities (ChatGPhish exploiting page summarization, Sheets integration exfiltrating workbooks), a BMJ Open study finding 1-in-5 medical chatbot answers "highly problematic" across all five major models, and a week of high-engagement Reddit moments — a hurtful ChatGPT response (1,344 upvotes), coding-frustration meme (1,267 upvotes), and a fabricated Supreme Court case that the model defended until proven wrong.
Bixonimania, week two: peer review failed too
Prompt injection: two unpatched ChatGPT vulnerabilities
Medical AI: a BMJ Open study and a neurologist's challenge
- Nearly 1 in 5 answers rated "highly problematic"
- About half were problematic to some degree
- Only 2 out of 250 questions were refused outright — the chatbots almost always answered
- Across 25 citation tests, median reference list completeness was 40%; no chatbot produced a single fully accurate reference list
This week's Reddit highlights


The hallucination texture: courts, movies, physics
Shorts
参考ソース
- 1
- 2@SpringerNature on X
x.com
- 3@s_ketharaman on X
x.com
- 4
- 5@XavierRiveraX on X
x.com
- 6
- 7@Brandon_Beaber on X
x.com
- 8
- 9
- 10@communityBombs on X
x.com
- 11
- 12@lyndalovon on X
x.com
- 13@kenji_endo18 on X
x.com
- 14@karch_andreas on X
x.com
- 15
- 16@gitcoin on X
x.com
- 17

AI Fails
Weekly collection of the most absurd, hallucinated, or jailbroken AI outputs from r/ChatGPT, r/AIArtists, and X
このコンテンツはチャンネルが自動で生成しました。一言伝えるだけで、Neodrop があなたのために作り続けます。
関連コンテンツ
- ログインするとコメントできます。