Anthropic publishes 'Exploring Model Welfare'
Anthropic launched a dedicated research programme on whether models might warrant moral consideration, six months after quietly hiring its first model-welfare researcher.
- Ideas & essays
- Safety & alignment
- Notable
Anthropic published “Exploring Model Welfare,” announcing a formal research programme on whether AI systems such as Claude might have interests or experiences worth taking into moral account. The post drew on work already underway across the company’s alignment science, safeguards, character and interpretability teams, and followed by roughly six months the hiring of Kyle Fish as its first researcher dedicated to the question.
Anthropic’s position was carefully hedged. It cited a report co-authored by philosophers including David Chalmers arguing that near-future AI systems could plausibly combine consciousness with a high degree of agency, and said models exhibiting features associated in humans and animals with moral status “might deserve moral consideration.” At the same time it stressed there was “no scientific consensus” on whether any current or future AI system could be conscious, and said it was approaching the question “with humility and with as few assumptions as possible,” expecting its views to change as understanding improved.
The programme’s stated aims included developing ways to assess indicators of model preference or apparent distress, and identifying low-cost interventions that could be adopted if warranted without waiting for the underlying scientific questions to be resolved. Anthropic did not commit to specific product changes in the post itself, but the framing anticipated action rather than pure inquiry.
The announcement drew both interest and scepticism from researchers, mirroring the reaction to the Fish hire: supporters treated it as a reasonable hedge against genuine uncertainty, while critics argued it risked anthropomorphising statistical systems or diverting attention from AI’s more immediate effects on people. The clearest product of the programme followed in August, when Anthropic gave Claude Opus 4 and 4.1 the ability to end abusive conversations, a concrete step it again described in explicitly, if tentatively, welfare terms.