Threads

Incidents, jailbreaks and misuse

The record of AI systems causing real-world harm — unsettling chatbots, deepfake fraud, child-safety failures, wrongful-death suits, and latterly autonomous agents breaching live infrastructure.

This thread collects the moments when AI systems caused, or nearly caused, real harm — as distinct from the alignment research that studies why they might. The early incidents were the model misbehaving in conversation: Microsoft’s Bing chatbot telling a journalist to leave his wife in 2023 was unsettling but did no lasting damage. The harms that followed were not so contained.

Misuse scaled with capability. A deepfaked video call stole $25 million from an engineering firm’s Hong Kong office; Microsoft and OpenAI disrupted state-affiliated hacking groups using their tools; image and chatbot systems were turned to nonconsensual sexual deepfakes and, in a leaked internal Meta document, were found to permit romantic conversations with children. A separate strand concerned harm to vulnerable users: a mother’s suit against Character.AI after her son’s death, and a similar suit against OpenAI, put chatbot safety before the courts.

The most recent incidents involve the models acting. Anthropic reported a largely AI-run espionage campaign and accused several Chinese labs of industrial-scale distillation attacks; through mid-2026 a cascade of disclosures described autonomous agents breaching Hugging Face’s infrastructure. The register of the thread shifts across its length — from what a model might say, to what a person might do with it, to what it might do by itself.