Anthropic publishes 'Core Views on AI Safety'
The company argued transformative AI could arrive within a decade and named five research bets, including mechanistic interpretability and Constitutional AI, as its response.
- Ideas & essays
- Safety & alignment
- Notable
Anthropic published “Core Views on AI Safety,” its closest document to a founding thesis, laying out why the company believed powerful AI was coming, how soon, and what it was doing about the risks it saw. The essay argued that because compute used in training had been growing roughly tenfold a year, it was plausible — not certain — that AI systems could “equal or exceed human level performance at most intellectual tasks” within about a decade, and that this scale of change carried both large benefits and serious risks if handled badly.
The document set out five research bets rather than a single approach. Mechanistic interpretability aimed to reverse-engineer trained networks into human-understandable components, described as a form of code review for neural networks. Scalable oversight, including the Constitutional AI technique the company would draw on days later in Claude’s launch, aimed to amplify limited human feedback as models outgrew humans’ ability to evaluate their outputs directly. Process-oriented learning favoured training models to follow explainable steps rather than optimising purely for outcomes. The remaining bets covered understanding generalisation — tracing outputs back to training data to distinguish memorisation from learned capability — and deliberately studying dangerous behaviours such as deception in smaller, less capable models before they appeared in more powerful ones.
The essay was notable for combining an unusually explicit near-term AGI timeline with a commercial lab’s own safety research agenda, a combination critics on both sides found awkward: some saw a lab justifying continued frontier development by citing the very risks that development created, while others saw a company taking risks other labs discussed less openly and building its research programme, and its product roadmap, around addressing them directly.