Person
Arun Jose
Commentary by Arun Jose
From the commentary rail — every link leaves the site for the original piece.
- 11 September 2026 · Redwood ResearchCoT controllability evals seem very under-elicitedSimple prompt optimizations can improve model capability to control their reasoning
- 18 June 2026 · Redwood ResearchThe distillation double bind: Distilling misaligned models either transfers misalignment or it doesn'tIf it transfers misalignment, we might get a misaligned model that’s easier to incriminate. If it doesn’t, we might get a capable benign replacement model.
- 18 May 2026 · Redwood ResearchIncriminating misaligned AI models via distillationSuppose we have a dangerous misaligned AI that can fool alignment audits, and distill it into a student model.
- 28 April 2026 · Redwood ResearchRecursive forecastingEliciting long-term forecasts from myopic fitness-seekers