Anthropic maps the 'Assistant Axis' persona vector across open models
An intervention called activation capping, which constrains a model's activations to normal range, cut harmful persona-drift responses by roughly half in testing.
Safety & alignment