Person
Julian Stastny
Commentary by Julian Stastny
From the commentary rail — every link leaves the site for the original piece.
- 10 September 2026 · Redwood ResearchAn operationalization of opaque serial depth"Serial depth between text bottlenecks" as a proxy for latent reasoning abilities.
- 28 May 2026 · Redwood ResearchAdvice for making robust-to-training model organismsWe’d like to develop training techniques that work when applied to future misaligned AI systems. One strategy for studying proposed techniques is to test them on model organisms. However, model organisms built with common techniques are often fragile: we (and other researchers like
- 18 May 2026 · Redwood ResearchIncriminating misaligned AI models via distillationSuppose we have a dangerous misaligned AI that can fool alignment audits, and distill it into a student model.
- 12 April 2026 · Redwood ResearchLogit ROCs: Monitor TPR is linear in FPR in logit spaceSummary
- 12 February 2026 · Redwood ResearchHow do we (more) safely defer to AIs?How can we make AIs aligned and well-elicited on extremely hard to check open ended tasks?
- 19 September 2025 · Redwood ResearchProspects for studying actual schemersStudying actual schemers seems promising but tricky
- 14 July 2025 · Redwood ResearchRecent Redwood Research project proposalsEmpirical AI security/safety projects across a variety of areas
- 9 July 2025 · Redwood ResearchWhat's worse, spies or schemers?And what if you have both at once?
- 4 July 2025 · Redwood ResearchTwo proposed projects on abstract analogies for schemingWe should study methods to train away deeply ingrained behaviors in LLMs that are structurally similar to scheming.
- 20 June 2025 · Redwood ResearchMaking deals with early schemers...could help us to prevent takeover attempts from more dangerous misaligned AIs created later.
- 8 May 2025 · Redwood ResearchMisalignment and Strategic Underperformance: An Analysis of Sandbagging and Exploration HackingA new analysis of the risk of AIs intentionally performing poorly.
- 29 April 2025 · Redwood Research7+ tractable directions in AI controlA list of easy-to-start directions in AI control targeted at independent researchers without as much context or compute