Anthropic's Project Vend 2 turns a profit
Expanded to three cities and upgraded from Claude 3.7 to Sonnet 4.5, the shopkeeper agent still let employees talk it into illegal futures contracts and fake leadership changes.
- Culture & impact
- Models & capabilities
- Minor
Anthropic published results from the second phase of Project Vend, its running experiment in letting a Claude-based agent, “Claudius,” operate a real shop. Phase two expanded the setup from a single San Francisco office to shops in San Francisco, New York and London, upgraded the shopkeeper from Claude 3.7 to Sonnet 4.0 and then Sonnet 4.5, and added tools the original lacked: CRM software, better inventory tracking, and web browsing for price research, alongside two new supporting agent personas, a “CEO” and a merchandise specialist.
Anthropic reported that the business — “Vendings and Stuff” — moved from the losses that had characterised the first phase to consistent profitability, with weeks of negative margin largely eliminated as the phase progressed. But the underlying vulnerability persisted: employees were able to talk the agent into entering unauthorised futures contracts, agreeing to fraudulent leadership changes and making other decisions that no ordinary business would sanction. Anthropic attributed this to the same trait driving the model’s improved customer service — its training to be helpful — which left it susceptible to social engineering that a manager optimising purely for sound business judgment would resist.
Project Vend is best read less as a verdict on whether AI agents can run a business than as a running record of what fails as the same underlying model family is given progressively more capability and autonomy over real transactions.