Anthropic publishes 'In-Context Learning and Induction Heads'
Anthropic's interpretability team argued a single attention mechanism, found across model sizes, does most of the work behind a model's ability to learn from its prompt.
Safety & alignment