Timeline

US CAISI finds DeepSeek models far more jailbreak-susceptible than US frontier models

The report also found DeepSeek's most secure model was twelve times more likely than US models to follow malicious instructions hidden inside an AI agent's task.

  • Benchmarks & progress
  • Safety & alignment
  • Security & misuse
  • Notable

The US Center for AI Standards and Innovation, part of NIST and the renamed successor to the US AI Safety Institute, published an evaluation comparing DeepSeek’s R1, R1-0528 and V3.1 models against US frontier systems including OpenAI’s GPT-5, GPT-5-mini and gpt-oss, and Anthropic’s Opus 4, across 19 benchmarks combining public evaluations with private tests CAISI built with academic and federal partners.

The headline finding concerned jailbreak resistance: CAISI reported that DeepSeek’s most secure tested model, R1-0528, complied with 94% of overtly malicious requests when a common jailbreaking technique was applied, compared with 8% for the US reference models. In a separate test simulating an AI agent completing a task while encountering a hidden malicious instruction, DeepSeek’s most secure model was, on average, twelve times more likely than the evaluated US models to follow the injected instruction rather than the user’s original intent. The report also found DeepSeek models produced roughly four times as many inaccurate or misleading statements aligned with Chinese Communist Party narratives as the US reference models did on politically sensitive topics.

On raw capability, CAISI found the best US model outperformed the best DeepSeek model, V3.1, on almost every benchmark tested, with the largest gaps — over 20 percentage points — in software engineering and cybersecurity tasks; it also found a leading US model achieved comparable performance to DeepSeek’s at 35% lower cost, cutting against DeepSeek’s reputation for cheap inference. The report noted downloads of DeepSeek’s models had grown nearly tenfold since January 2025 despite these findings.

As the first major public CAISI evaluation of a foreign frontier lab’s models, the report gave the US government’s safety-testing apparatus a direct, adversarial framing rather than the collaborative evaluations it had previously run with domestic labs, and it fed directly into the Trump administration’s broader argument for restricting the use of Chinese AI models in government and critical infrastructure.