Claude 3.5 Sonnet and Artifacts change how people use chatbots
Priced and sped like Anthropic's mid-tier model, it scored 64% on the company's internal agentic-coding evaluation against 38% for the outgoing flagship.
- Models & capabilities
- Major
Anthropic released Claude 3.5 Sonnet, a mid-tier model that it said outperformed its own flagship, Claude 3 Opus, released three months earlier, on a wide range of evaluations, while running at twice Opus’s speed and at Sonnet-tier pricing — $3 per million input tokens and $15 per million output tokens, against Opus’s substantially higher rate. On Anthropic’s internal agentic coding evaluation, Claude 3.5 Sonnet solved 64% of problems, compared with 38% for Claude 3 Opus, and the company reported gains on graduate-level reasoning (GPQA), undergraduate knowledge (MMLU) and code generation (HumanEval) benchmarks as well. That a model priced and marketed as the middle tier of Anthropic’s line-up beat the company’s own top-priced model was a first for the Claude family and briefly reframed “flagship” as a claim about the newest release rather than the most expensive one.
Alongside the model, Anthropic launched Artifacts, a side panel in Claude.ai that displayed and let users interact with substantial pieces of content the model generated — code, documents, website mock-ups, diagrams — separately from the conversation thread, updating live as the user asked for revisions. Anthropic described it as marking Claude’s shift “from a conversational AI to a collaborative work environment,” an explicit move to compete on workflow rather than on raw chat quality alone.
The release landed in the middle of an intensifying contest between OpenAI, Google and Anthropic over which company’s assistant developers would build on top of by default, a contest that increasingly turned on coding ability and integration into existing tools rather than benchmark scores alone. Claude 3.5 Sonnet’s coding performance and Artifacts’ code-preview workflow contributed to Anthropic’s growing reputation among developers over the following year, a reputation that would shape the company’s later push into dedicated coding products.