DeepMind publishes cybersecurity evaluation framework for frontier AI
The framework scored models against 50 challenges spanning the whole attack chain, drawing on more than 12,000 real attempts to misuse AI for cyberattacks across 20 countries.
- Safety & alignment
- Security & misuse
- Minor
Google DeepMind published a framework for evaluating how advanced AI models could be misused to mount cyberattacks, assembled by researchers including Four Flynn, Mikel Rodriguez and Raluca Ada Popa. The company described it as the most comprehensive evaluation of its kind published to that point.
The framework broke a cyberattack into distinct phases — reconnaissance, weaponisation, delivery, exploitation and action on objectives, among others — and built 50 challenges spanning that full chain, covering seven categories of attack including phishing, malware development and denial-of-service. To ground the evaluation in real-world behaviour rather than hypothetical threats, the researchers drew on Google’s Threat Intelligence Group data on more than 12,000 documented attempts to use AI in cyberattacks across more than 20 countries, using this to identify which stages of an attack AI assistance most changed the economics of — highlighted as evasion and persistence, phases that had received comparatively little attention from prior cyber-capability evaluations focused on exploit generation.
DeepMind’s own assessment, reported alongside the framework, was that current models were “unlikely to enable breakthrough capabilities” for attackers when used in isolation, but that the overlooked bottleneck stages were where future gains in offensive usefulness were most likely to appear first.
The publication extended a pattern set by other labs’ dangerous-capability evaluations — OpenAI and Anthropic had published comparable cyber-offence assessments in their own model system cards — but positioned itself as a general-purpose framework other evaluators could apply, rather than a one-off assessment of a single model.