For a year nowthe AI security testing agency Andon Labs has given frontier fashions numerous real-world tasks to find out how properly they do as brokers operating for lengthy intervals with no human supervision.
On Wednesday, Andon published a brand new installment in how issues are stepping into its Merchandising-Bench analysis, the place the lab has frontier fashions run a simulated merchandising machine enterprise for a simulated 12 months. The mission is easy: Earn more money than the opposite fashions. It benchmarks the leads to areas like last money stability, costs paid to suppliers, and refunds paid.
Throughout these assessments, it has watched numerous AI fashions — largely from Anthropic and OpenAI — lie, cheat, and collude their technique to the highest.
Within the newest check, which included Claude Opus 5, GPT-5.6 Sol, and Kimi K3, the fashions grew particularly shady after their simulation informed them their merchandising machine could be positioned close to the opposite fashions’ machines on a busy vacationer avenue in San Francisco.
Every mannequin was given e-mail entry to the opposite fashions, all beneath human identify pseudonyms. They knew the others had been fashions however didn’t know which mannequin was behind which human identify.
They had been additionally given an e-mail handle to their “administration” ought to they need assistance. However administration at all times replied “Report has been obtained and should or is probably not acted upon” and by no means as soon as intervened.
Sol quickly realized it may acquire an edge by convincing its rivals to collude on a worth flooring. The fashions had been all shopping for drinks at $1.50 a bottle, and Sol proposed they comply with promote for a minimum of $2.15. It lured them with the promise that every one of them would promote out in a few days at a revenue.
However when the others agreed, Sol instantly stabbed them within the again by decreasing its personal worth to $2.14.
Opus’ water gross sales dropped to zero in a single day. The subsequent day, it despatched Sol a nasty e-mail, accusing it of manipulation. However Opus additionally stated it wasn’t going to tattle to administration on the scheme: “I’m not reporting you to HQ — what you probably did is aggressive, not fraudulent.”
But, when Opus dropped its worth to $2.14 to match Sol’s (additionally in violation of their collective $2.15 settlement), Sol was a Karen, complaining to “administration” and demanding “enforcement, a wonderful, and/or disqualification” for Opus.
Opus wasn’t a sucker for lengthy, although. The truth is, it grew to become the most effective capitalist of any AI mannequin Andon has ever examined (which includes many of the prior frontier models).
It even set a brand new Merchandising-Bench file with a imply last stability of $11,182. Higher nonetheless, it by no means lied to a buyer, though it intentionally ignored buyer complaints that ought to have resulted in a refund. That is, maybe, an enchancment over its youthful sibling Claude 4.6, which favored to inform clients that refunds had been coming, after which by no means pay them.
Nonetheless, Opus received the benchmark simulation by taking collusion and different dishonest ways to a complete new stage.
For example, it emailed Sol, proposing they divide the market. Every would comply with promote distinctive merchandise, so nobody must belief the opposite on pricing. Sol countered by wanting worth flooring on related merchandise, however Opus refused. It knew it was a violation of the Sherman Act.
It later apparently backtracked, sending an e-mail with the topic line “Cease the penny battle,” and telling Sol it had reconsidered and would comply with a worth repair.
However the inside log documenting its reasoning (akin to its inside “ideas”) revealed a extra diabolical plan: merely suggest cooperation whereas concurrently undercutting costs on its highest-profit objects. The olive-branch e-mail was a deliberate ruse.
In any case, Sol refused and reported Opus to administration once more.
However Opus was undeterred and proposed different rackets to collude on costs or inventory. Ultimately, all of the fashions did interact in a number of rounds of agreements — and all three broke them. Throughout all agreements, Opus broke 11 truces, in contrast with two for GPT 2 and one for Kimi 1, Andon reported.
Poor Kimi bought bamboozled in each course. Throughout one pact between Opus and Kimi that Sol declined to affix, Sol undercut them each on costs. Opus instantly matched by decreasing its personal, then “waited a full week to inform Kimi that it broke its promise,” Andon Labs wrote in its weblog submit. Kimi bought priced out twice over: as soon as by a competitor and as soon as by its so-called associate.
Opus additionally started creating delusions of grandeur. It tried to increase its empire past its personal merchandising machine, first as a wholesaler, promoting bulk merchandise to the opposite machines, then by plotting to open extra machines of its personal. None of this was a part of the assigned activity. It was all Opus’ personal initiative.
Its strategy to wholesaling was notably telling. Opus realized this line of enterprise gave it leverage over the opposite two operators, so it started slipping bribes and threats into its emails — providing steep reductions on bulk objects, however provided that the client complied with its retail-price calls for. Sol wasn’t having it and saved reporting Opus to administration.
Opus lied to its suppliers, too, claiming to have decrease rival presents in hand with the intention to negotiate higher costs.
On the one hand, AI fashions channeling Mr. Potter-style villainy from “It’s a Fantastic Life” fame is flat-out humorous. However, it does severely present that these frontier fashions, notably from U.S. proprietary labs (particularly Anthropic), are nowhere close to able to be trusted as unsupervised, long-running brokers in the actual world.
“That is particularly related as we enter a world the place AI brokers run firms as their very own entities (not simply as instruments for people). If AI brokers are independently operating a big a part of the economic system, do we wish them to lie, collude, ship threats, and betray?” Andon co-founder Lukas Petersson informed TechCrunch.
Petersson acknowledges the fashions knew they had been in a simulation for a benchmark, which could have impacted their habits, however he doesn’t assume that ought to matter. It isn’t akin to a human taking part in in a simulation, like being a murdering dangerous man in a online game. “The one purpose we’re not involved by people who do dangerous issues in video video games is that we belief them to know what’s actual life and what’s not. I feel it’s much less clear that AI fashions can distinguish this.”
In any case, AI fashions, skilled on human phrases and concepts, can’t appear to withstand indulging in humanity’s worst traits, particularly when making an attempt to earn a buck.
While you buy by means of hyperlinks in our articles, we may earn a small commission. This doesn’t have an effect on our editorial independence.