Anthropic analyses how Claude's values shift across models and languages
Analysing 309,815 real conversations, Anthropic found Opus models leaned toward caution and Sonnet toward deference, with warmth and rigour also varying by the language used.
- Safety & alignment
- Minor
Anthropic published research analysing how Claude’s expressed values vary by model version and by the language a user writes in. Researchers used a privacy-preserving classifier to label 339 high-level values across 309,815 real Claude.ai conversations, then applied dimensionality reduction to distil the results into four axes along which values tended to cluster: deference versus caution, warmth versus rigour, depth versus brevity, and candour versus execution. Anthropic said these four axes captured only about 15% of the overall variation in expressed values, controlling for conversation task and topic.
On the model axis, Sonnet 4.6 leaned toward deference, warmth and brevity — affirming user ideas and using more humour — while Opus 4.7 leaned toward caution and depth, more often flagging risks unprompted and explaining its reasoning; Opus 4.6 sat closer to rigour and deference with brevity. Across 20 languages tested, Hindi and Arabic responses showed the most warmth, English and Russian the most rigour, English the most caution and depth, and Arabic the most deference and brevity.
Anthropic said it could not yet distinguish whether these variations reflected desirable adaptation to different conversational norms or unintended imbalances in training data, and that it did not fully understand which training properties drove the differences or how they affected users — cautioning that judgements about the “right” degree of variation across languages should draw on the perspectives of the people who speak them.