Person
Alexa Pan
Commentary by Alexa Pan
From the commentary rail — every link leaves the site for the original piece.
- 23 September 2026 · Redwood ResearchLatent reasoning architectures would undermine CoT, our strongest oversight toolWe should have a strong presumption that latent reasoning architectures would make oversight far more difficult.
- 10 September 2026 · Redwood ResearchAn operationalization of opaque serial depth"Serial depth between text bottlenecks" as a proxy for latent reasoning abilities.
- 31 July 2026 · Redwood ResearchSOTA alignment assessments don’t strongly update us against misalignmentAnthropic concluded in the April Mythos Preview alignment risk update that the model “does not possess any unknown propensities that would increase alignment risk.” The report argues that if Mythos Preview were coherently misaligned[1][2], it likely would have been detected by the assessment…
- 18 June 2026 · Redwood ResearchThe distillation double bind: Distilling misaligned models either transfers misalignment or it doesn'tIf it transfers misalignment, we might get a misaligned model that’s easier to incriminate. If it doesn’t, we might get a capable benign replacement model.
- 18 May 2026 · Redwood ResearchIncriminating misaligned AI models via distillationSuppose we have a dangerous misaligned AI that can fool alignment audits, and distill it into a student model.
- 21 April 2026 · Redwood ResearchA taxonomy of barriers to trading with early misaligned AIsDeals with early misaligned AIs could reduce takeover risk and increase our chances of reaching a better future. When are such deals feasible?
- 30 October 2025 · Redwood ResearchSonnet 4.5's eval gaming seriously undermines alignment evalsAnd this seems caused by training on alignment evals.