Timeline

Anthropic previews Claude Mythos, withheld from public release over cyber-offense capability

Anthropic reported the model wrote a working Firefox exploit in 181 of several hundred attempts, versus two for its predecessor Opus 4.6, and found a 27-year-old OpenBSD bug.

  • Safety & alignment
  • Security & misuse
  • Models & capabilities
  • Major

Anthropic disclosed Claude Mythos Preview, describing it as its most capable model yet for coding and agentic tasks, and said it had no plan to release it to the public because of the cyberattack capability it demonstrated. The model’s existence had already leaked ten days earlier, when security researchers found roughly 3,000 unpublished Anthropic assets, including a draft blog post naming the model, in an unsecured database.

Anthropic’s own account of the model’s cybersecurity behaviour was unusually specific. Mythos identified zero-day vulnerabilities across “every major operating system and every major web browser” when directed to, including a 27-year-old bug in OpenBSD’s SACK handling and a flaw in FFmpeg that had survived roughly five million automated fuzzing runs. It then developed exploits autonomously and without human steering: a Linux privilege-escalation chain exploiting race conditions and KASLR bypasses, and a remote-code-execution exploit against FreeBSD’s NFS server built from a 20-gadget ROP chain split across multiple network packets. Anthropic said the model produced a working Firefox sandbox-escape exploit in 181 attempts out of several hundred, compared with two successes for its predecessor, Claude Opus 4.6, over the same number of attempts — and reported a score of 83.1% on CyberGym’s vulnerability-reproduction benchmark against 66.6% for Opus 4.6.

Rather than release the model broadly, Anthropic said it would make it available only to a limited set of partner organisations through a defensive programme, Project Glasswing, and would explore public access later behind new safeguards. Commentators including security researcher Bruce Schneier noted the company was, in effect, asking the world to trust a single lab’s judgement about when an offensive capability was safe to withhold versus share, with no independent verification of the claims. Anthropic argued the alternative — waiting for attackers to reach the same capability first — was worse.

The disclosure was widely described as the first case since GPT-2 in 2019 of a lab publicly withholding a model on capability grounds, though the comparison was inexact: GPT-2 was withheld over misuse of text generation, Mythos over vulnerability discovery and exploit-writing specifically. The following months saw Project Glasswing’s partner list and reported vulnerability counts expand substantially, and set a precedent other labs faced pressure to match.