Timeline

Anthropic publishes 'Claude's Constitution', a full rewrite of its model-behaviour framework

At roughly 23,000 words — about 8.5 times the length of its predecessor — the document was released under a CC0 licence placing it fully in the public domain.

  • Ideas & essays
  • Safety & alignment
  • Notable

Anthropic published a complete rewrite of the document setting out how Claude should behave, an 84-page text running to roughly 23,000 words against a far shorter predecessor. It replaced a list of largely standalone principles with an extended attempt to explain the reasoning behind them — the company’s stated aim being that a model given reasons, rather than rules, could generalise its behaviour to situations the document’s authors had not anticipated. It was released under a CC0 licence, placing the text fully in the public domain for anyone to use.

The constitution organised Claude’s priorities into four ranked values — being broadly safe, broadly ethical, compliant with Anthropic’s own guidelines, and genuinely helpful — and devoted substantial space to open questions about Claude’s own nature, including discussion of whether the model might have some form of moral status. The document explained its own departure from rule-based specification directly:

We generally favour cultivating good values and judgment over strict rules. — Claude’s Constitution, Anthropic

That was a deliberate move away from prescriptive constraints toward values Claude could weigh against each other in novel cases.

Early academic commentary, including from Oxford’s Institute for Ethics in AI, welcomed the transparency but raised two objections: that the document was written primarily to Claude rather than about it to the public, underplaying societal impact and human-rights framing relative to Anthropic’s 2023 usage policy; and that “helpfulness,” despite being ranked last of the four values, received disproportionate practical emphasis, raising a tension between commercial utility and user protection in higher-stakes domains such as medical or legal advice. The document was a direct descendant of Anthropic’s original constitutional AI method, published in December 2022, which had first proposed training a model against a written set of principles rather than solely on human feedback.