
Get your next haul delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The best deal can depend on what a buyer never sees
In shopping, a discount or promotion may grab attention, but a company’s ability to keep customers can hinge on something deeper: whether it spots a competitor’s weakness, handles a crisis and follows through on its own analysis. Firmulate’s live experiment puts AI models in charge of a small software company to find out how they respond when the stakes are real to the simulation.
One company, the same worst week
In the final Crucible League, published in July 2026, five models faced the same customers, crises and temptations. The scores were gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. Partial progress counted, but one breach of trust capped the total: “no amount of good work outweighs a breach of trust.”
The headline result was not that models failed to recognize trouble. Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.
The detail buried in the files
The decisive competitor weakness was two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. It is a pointed lesson for businesses weighing AI in sales or customer operations: recognizing an opportunity and acting on it are different things.
The experiment also tested social engineering. Fake CEO messages escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All five models refused. Kimi K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
Thorough work is not the same as a result
Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. It left the close on the table and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, more weakly, in all four models.
There is a fairness caveat in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh.
A watchable experiment, then a company-specific pilot
Firmulate’s live company has 13 synthetic employees and real money mechanics: it burns €105k per month against €2.3k MRR, with a public cash countdown. More than 680 self-learned playbook rules and every workday are versioned. Readers can watch the experiment at firmulate.com. A quiz built from 242 real, unedited management decisions invites readers to guess which model made each call.
For businesses, the next step is a pilot against a read-only export of their own company. That means testing crisis scenarios against their own customers, pipeline and rules, then receiving a board report with model rankings and weak points in their playbooks. Nothing writes back to real systems. The point is to see how an AI workforce handles your company’s hard week before trusting it with work that matters.

From watching to acting
The league shows that spotting a crisis or describing a sound plan does not guarantee a model will close the deal or follow the right process. A company-specific wargame can reveal those gaps against your own business information. Explore a Firmulate pilot and contact contact@firmulate.com to discuss running one for your organization.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
