DeepMind's Sparrow explores rule-based RLHF for safer dialogue
Adversarial testers broke Sparrow's written safety rules in about 8% of attempts, roughly a third of the rate for a baseline model tested the same way.
Models & capabilities · Safety & alignment