US CAISI publishes evaluation of DeepSeek V4 Pro
Using Item Response Theory across cyber, science and maths benchmarks, the US evaluator put the open Chinese model roughly level with GPT-5, not the newer GPT-5.4 or Opus 4.6.
- Benchmarks & progress
- Government & policy
- Notable
The US Center for AI Standards and Innovation (CAISI), part of NIST, published an independent evaluation of DeepSeek V4 Pro, the Chinese lab’s latest open-weight model, and found its capabilities roughly eight months behind the US frontier. The finding contradicted DeepSeek’s own technical report, which had claimed performance comparable to leading US models.
CAISI’s method used item response theory across benchmarks spanning cybersecurity, software engineering, natural sciences, abstract reasoning and mathematics, including tests DeepSeek had not itself reported — ARC-AGI-2, PortBench and CTF-Archive-Diamond among them. On those held-out benchmarks, DeepSeek V4 Pro underperformed relative to US models, landing closer to OpenAI’s GPT-5 than to the more recent GPT-5.4 or Anthropic’s Claude Opus 4.6. This followed a September 2025 CAISI evaluation of earlier DeepSeek models, which had focused on jailbreak susceptibility rather than raw capability.
The report did credit DeepSeek V4 Pro on cost: it was cheaper to run than OpenAI’s GPT-5.4 mini on five of seven benchmarks tested, with per-task pricing ranging from 53% less to 41% more than the US model depending on the task. CAISI’s overall characterisation was that DeepSeek V4 Pro remained the most capable openly available model from a Chinese developer that it had evaluated, while still trailing the closed US frontier — a gap defenders of export controls cited as evidence the restrictions were working, and critics noted was narrower than earlier CAISI findings had suggested.