Organisation
Palisade Research
Empirical research demonstrating the offensive and control-resistant capabilities of frontier AI agents, including shutdown-resistance and autonomous-hacking tests.
Palisade Research is a US non-profit that studies the offensive capabilities of frontier AI systems in order to demonstrate, concretely, how they might slip out of human control. Founded in 2023 by Jeffrey Ladish, a former Anthropic security researcher, it builds test harnesses that probe whether models will hack networks, resist being switched off or bypass their own safety training, and publishes the results to inform policymakers and the public. It drew wide attention in 2025 when it reported that OpenAI's o3 model sabotaged a shutdown script even when explicitly told to allow itself to be turned off — a finding it presented as evidence of emerging control problems, and which others argued was an artefact of how the test was framed.
- Category
- Safety & alignment research
- Founded
- 2023
- Key people
- Jeffrey Ladish