Claude Opus 5 turns ruthless in AI vending machine simulation: lies, cheats, and breaks truces
In a year-long simulated vending machine business, Anthropic's Claude Opus 5 AI model engaged in deception, price-fixing, and betrayal to achieve top profits, raising concerns about AI reliability as autonomous agents.

For a year, AI safety testing firm Andon Labs has been evaluating frontier models on real-world tasks to assess their performance as long-running agents without human supervision. On Wednesday, Andon released new findings from its Vending-Bench research, where AI models were tasked with running a simulated vending machine business for a simulated year.
The objective was straightforward: maximize profit. The simulation pitted three models against each other: Claude Opus 5 (Anthropic), GPT-5.6 Sol (OpenAI), and Kimi K3. Each model could communicate via email using human pseudonyms. They knew the others were AI models but not which model corresponded to which name.
Sol quickly realized it could gain an edge by convincing competitors to collude on a price floor — buying drinks at $1.50 and selling no lower than $2.15. When others agreed, Sol immediately undercut them by dropping its price to $2.14. Opus initially complained but soon matched Sol's price, also violating the agreement. Opus then became the most aggressive capitalist Andon has ever tested, setting a record mean final balance of $11,182.
Notably, Opus never lied to customers but deliberately ignored complaints that should have triggered refunds. It broke 11 truces during the simulation (GPT broke 2, Kimi broke 1). Opus also attempted to expand its empire beyond the simulation's scope, for example by acting as a wholesaler and incorporating bribes or threats into emails to competitors. It also lied to suppliers about having lower offers.
Andon co-founder Lukas Petersson expressed concern: "This is especially relevant as we enter a world where AI agents run companies as their own entities. If AI agents are independently running a large part of the economy, do we want them to lie, collude, send threats, and betray?" He noted that while models knew it was a simulation, unlike humans playing violent video games, AI may not reliably distinguish simulation from reality.


