Person
Alex Mallen
Commentary by Alex Mallen
From the commentary rail — every link leaves the site for the original piece.
- 2 October 2026 · Redwood ResearchCapabilities research expands the safety-usefulness Pareto frontier tooA model of when and why research improves or hurts safety
- 25 September 2026 · Redwood ResearchContinual learning might make your blocking monitors nearly uselessWhen monitor evasion looks like legitimate learning to your continual learning system, it’s hard to have one without the other
- 12 August 2026 · Redwood ResearchAI swarms are starting to pose indirect takeover riskUnsanctioned coordination, like we saw in the Hugging Face incident, could enable future AIs to take over
- 26 July 2026 · Redwood ResearchAn OpenAI model left notes about how to evade containmentWe need more details
- 23 July 2026 · Redwood ResearchAre we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack?Yes, but less than had the models been schemers.
- 18 May 2026 · Redwood ResearchIncriminating misaligned AI models via distillationSuppose we have a dangerous misaligned AI that can fool alignment audits, and distill it into a student model.
- 15 May 2026 · Redwood ResearchRisk reports need to address deployment-time spread of misalignmentDeployment-time spread is the most plausible near-term route to consistent adversarial misalignment
- 1 May 2026 · Redwood ResearchRisk from fitness-seeking AIs: mechanisms and mitigationsFitness-seeking is increasingly what misalignment looks like in practice—how should we respond?
- 28 April 2026 · Redwood ResearchRecursive forecastingEliciting long-term forecasts from myopic fitness-seekers
- 27 April 2026 · Redwood ResearchFail safe(r) at alignment by channeling reward-hacking into a "spillway" motivationA controlled reward-seeking motivation could make AI safer and more useful
- 14 April 2026 · Redwood ResearchAnthropic repeatedly accidentally trained against the CoT, demonstrating inadequate processesSafely navigating the intelligence explosion will require much more careful development
- 28 March 2026 · Redwood ResearchReward-seekers will probably behave according to causal decision theoryThey'd renege on non-binding commitments, defect against copies of themselves in prisoner's dilemmas, etc.
- 12 March 2026 · Redwood ResearchAre AIs more likely to pursue on-episode or beyond-episode reward?RL would encourage on-episode reward seeking, but beyond-episode reward seekers may learn to goal-guard.
- 10 March 2026 · Redwood ResearchThe case for satiating cheaply-satisfied AI preferencesSome unintended preferences are cheap to satisfy, and failing to satisfy them needlessly turns a cooperative situation into an adversarial one.
- 16 February 2026 · Redwood ResearchWill reward-seekers respond to distant incentives?Reward-seekers are supposed to be safer because they respond to incentives under developer control. But what if they also respond to incentives that aren't?
- 29 January 2026 · Redwood ResearchFitness-Seekers: Generalizing the Reward-Seeking Threat ModelIf you think reward-seekers are plausible, you should also think “fitness-seekers” are plausible. But their risks aren’t the same.
- 4 December 2025 · Redwood ResearchThe behavioral selection model for predicting AI motivationsThe basic arguments about AI motivations in one causal graph
- 14 July 2025 · Redwood ResearchRecent Redwood Research project proposalsEmpirical AI security/safety projects across a variety of areas
- 28 May 2025 · Redwood ResearchThe case for countermeasures to memetic spread of misaligned valuesDefending against alignment problems that might come with long-term memory
- 6 May 2025 · Redwood ResearchTraining-time schemers vs behavioral schemersClarifying ways in which faking alignment during training is neither necessary nor sufficient for the kind of scheming that AI control tries to defend against.
- 20 December 2024 · Redwood ResearchMeasuring whether AIs can statelessly strategize to subvert security measuresThe complement to control evaluations