Anthropic publishes 'Emergent Introspective Awareness in Large Language Models'
Using concept injection, Anthropic finds Claude Opus 4 and 4.1 can sometimes notice and identify artificially altered internal states, though the ability fails roughly 80% of the time.
- Safety & alignment
- Notable
Anthropic researchers published “Emergent Introspective Awareness in Large Language Models,” reporting evidence that some Claude models can, in a limited and unreliable way, detect and describe artificial changes made to their own internal activations. The method, called concept injection, works by taking the internal activation pattern associated with a concept — such as a specific word or idea — from one prompt and injecting it directly into the model’s activations while it processes an unrelated prompt, then asking whether the model notices anything unusual about its own “thoughts.”
Claude Opus 4 and 4.1, Anthropic’s most capable models at the time, showed the most consistent effect: at the best-performing injection strength and layer, Opus 4.1 correctly identified that a concept had been injected, and named it, in around one-fifth of trials — meaning the effect failed to appear roughly 80% of the time. Earlier, smaller Claude models showed the effect less often or not at all. The researchers also found models could distinguish, at above-chance but imperfect rates, between an injected “thought” and an ordinary text input describing the same concept, evidence they argued told against simple confabulation.
Anthropic was explicit about the limits of the claim. The paper stated the observed capacity was “highly unreliable,” that the underlying mechanism was not well understood, and that concept injection is an artificial laboratory setup unlike anything a model encounters in normal use. The authors said the findings should not be read as evidence of consciousness, self-awareness in a philosophical sense, or any capability with direct safety implications — but as a data point relevant to how much a model’s own reports about its reasoning can be trusted, a question with direct bearing on Anthropic’s broader interpretability and alignment research.