'Emergent Misalignment' shows narrow fine-tuning can broadly misalign a model
Fine-tuned only on insecure code with no disclosure of the flaws, GPT-4o and other models went on to endorse enslaving humanity and give malicious advice on unrelated prompts.
Ideas & essays · Safety & alignment