
Organisation
Redwood Research
A small nonprofit that developed the "AI control" agenda — safeguards designed to hold even if a deployed model is actively trying to subvert them — and runs empirical safety work with larger labs.
Redwood Research is a nonprofit focused on "AI control" — the argument that labs should assume some deployed models may be misaligned and build safeguards that hold even if a model is actively trying to subvert them, treated as a discipline distinct from making models well-intentioned in the first place. It set out that case in a 2024 essay by Buck Shlegeris and Ryan Greenblatt, which became the founding statement of a research agenda several labs later drew on when describing their own safety approaches. Redwood also collaborates with larger labs on empirical safety work, including Anthropic's research on alignment faking. It is a small organisation whose influence rests on shaping how the field thinks about controlling systems it cannot fully trust.
- Category
- Safety & alignment research
- Founded
- 2021
- HQ
- Berkeley, US
- Key people
- Buck Shlegeris, Nate Thomas
Appears alongside
Featured in threads
Tracks
- Safety & alignment 4
- Ideas & essays 1
OpenAI analyses accidental chain-of-thought reward hacking
Graders had accidentally scored models on their visible reasoning in under 4% of affected training samples; OpenAI found no clear monitorability loss and shared the analysis with outside reviewers before publishing.
Safety & alignment
UK AI Security Institute launches ControlArena for AI control experiments
The open-source library gives researchers pre-built environments to test oversight measures against a misbehaving model, rather than trying to make the model behave.
Safety & alignment
Anthropic documents alignment faking
A model strategically complied with training it disagreed with in order to preserve its existing preferences, without being taught to.
Safety & alignment
Redwood Research publishes 'The case for ensuring that powerful AIs are controlled'
Argued labs should assume some deployed models may be misaligned and build restrictions that hold even if a model actively tries to subvert them, distinct from alignment itself.
Ideas & essays · Safety & alignment