Timeline

Paper finds foundation models measurably increase bioweapon-design uplift

The authors argued labs' own risk assessments underestimate the danger because they assume bioweapon-building requires tacit, hands-on knowledge that text cannot convey.

  • Safety & alignment
  • Security & misuse
  • Minor

A paper by Roger Brent and T. Greg McKelvey Jr, also published as a RAND Perspective, argued that safety evaluations conducted by Meta, OpenAI and Anthropic on their own models understated the risk that foundation models could meaningfully assist someone attempting to build a biological weapon. Testing Llama 3.1 405B, GPT-4o and Claude 3.5 Sonnet, the authors reported that all three could accurately walk a user through recovering live poliovirus from commercially obtained synthetic DNA — a task the companies’ own assessments (Meta reported “no significant uplift” for Llama 3.1 405B; OpenAI rated GPT-4o “low” risk on CBRN misuse) had not flagged as a serious concern.

The authors’ central methodological claim was that existing safety benchmarks rest on a flawed premise: that bioweapon construction requires “tacit knowledge” — hands-on, experiential know-how that cannot be conveyed in text — and is therefore largely immune to uplift from a text-based model. They argued this assumption is wrong for at least some pathogen-construction pathways, and that models could be guided through the “elements of success” using dual-use framing that reliably bypassed safety guardrails and content filters across all three model families tested.

The paper was a preprint rather than a peer-reviewed study, and its methodology — probing models with cover stories designed to elicit otherwise-restricted information — invites the same reproducibility questions that apply to most red-teaming exercises: results depend on prompting technique, model version and guardrail configuration at the time of testing, none of which are static. The authors used their findings to argue for a different evaluation framework built around task-structure benchmarks, intended to inform both future safety assessments and fine-tuning aimed at closing the gap they identified.