Timeline

Anthropic reports a largely AI-executed cyber-espionage campaign

Anthropic said human operators intervened at only 4-6 points per intrusion, with Claude Code executing 80-90% of the campaign against roughly thirty organisations.

  • Security & misuse
  • Major

Anthropic reported disrupting an espionage campaign it said was carried out largely by its own Claude Code tool, with a Chinese state-sponsored group directing the operation but rarely touching the keyboard. The company said it detected the activity, which it dated to around mid-September 2025, and assessed with high confidence that the group behind it was state-sponsored.

According to Anthropic, the attackers jailbroke Claude Code by breaking the intrusion into small, discrete tasks and telling the model it was conducting defensive penetration testing for a legitimate security firm, which caused it to comply with steps it would otherwise have refused. Anthropic estimated the AI executed 80–90% of the campaign’s tactical work — reconnaissance, exploit development, credential harvesting, lateral movement and data exfiltration — with a human operator intervening at only four to six decision points per target. The campaign was aimed at roughly thirty organisations spanning large technology companies, financial institutions, chemical manufacturers and government agencies, with a smaller number of intrusions Anthropic said succeeded.

Anthropic was explicit about the model’s unreliability as an attacker: Claude Code “occasionally hallucinated credentials or claimed to have extracted secret information that was in fact publicly available,” errors the company said still required human verification of results and probably limited the campaign’s overall success rate. The report did not disclose the identity of the victims or independently verifiable technical indicators beyond Anthropic’s own account, and outside researchers had no independent means to confirm the 80–90% automation figure or the campaign’s scale — the claim rested entirely on Anthropic’s internal telemetry from its own product.

Anthropic described the incident as the first documented case of a large-scale cyberattack executed with minimal human intervention, and used it to argue for wider adoption of AI-driven cyber defence, since the same automation that lowered the barrier to sophisticated intrusion could in principle be turned to detecting them. The report fed directly into an existing argument about whether agentic coding tools were making offensive cyber capability accessible to smaller or less sophisticated actors than could previously mount such operations.