Redwood Research publishes 'The case for ensuring that powerful AIs are controlled'
Argued labs should assume some deployed models may be misaligned and build restrictions that hold even if a model actively tries to subvert them, distinct from alignment itself.
Ideas & essays · Safety & alignment