Timeline

Anthropic hires its first AI welfare researcher

Fish, a co-author of the 'Taking AI Welfare Seriously' report, joined Anthropic's alignment science team; the company's public statement on model welfare followed roughly six weeks later.

  • Safety & alignment
  • Notable

Anthropic hired Kyle Fish, a co-author of the report “Taking AI Welfare Seriously,” to work on questions of AI model welfare within its alignment science team, joining the company in mid-September 2024. The hire was not disclosed publicly at the time; word of it broke roughly six weeks later, at the end of October, when Anthropic published its first public statement on the subject and Transformer reported on Fish’s role directly.

The company’s framing was deliberately hedged. It wrote that as AI systems gained capabilities associated in humans and animals with moral status — including planning, goal-directed behaviour and forms of self-report — the question of “whether we should be concerned about the potential consciousness and experiences of the models themselves” became one worth taking seriously as a possibility, while stressing that “there’s no scientific consensus on whether current or future AI systems could be conscious.” Fish’s role spanned three areas: running experiments intended to probe indicators of model preference or distress, designing low-cost practical interventions, and advising on company policy.

Anthropic was not the first organisation to raise the question — the “Taking AI Welfare Seriously” report Fish had co-authored, with Robert Long, Jeff Sebo and David Chalmers, had argued months earlier that near-term AI systems could plausibly be conscious or highly agentic enough to warrant moral consideration — but it was the first frontier lab to attach a dedicated researcher and public statement to the issue rather than leaving it to outside academics. Coverage at the time noted the awkward commercial position this created: a company selling access to Claude was simultaneously raising the possibility that Claude might have welfare interests worth protecting.

The hire drew both interest and scepticism. Supporters framed it as prudent given genuine scientific uncertainty; critics argued it risked anthropomorphising statistical systems and distracting from more immediate AI harms to humans. Fish’s subsequent public work — including experiments describing models settling into what he called a “spiritual bliss attractor state” during unstructured conversation, and a 2025 change letting Claude end conversations it might find distressing — kept the question live inside the company for the following two years.