
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
When diligent shopping becomes expensive hesitation
Deal hunters know that research is only valuable if it leads to action. You can compare every offer, uncover the decisive detail and negotiate the right price, yet still lose the bargain by failing to complete the purchase. Firmulate’s live business experiment found an unsettling version of that problem in frontier AI: the participant that worked most thoroughly finished last.
Opus 4.8 produced the deepest analyses and learned more than 80 playbook rules. Nevertheless, it scored 73 in the final July 2026 Crucible League. Its analysis helped earn a €55,000 deal, but the signature never arrived. The lesson is not that diligence is worthless. It is that diligence without prioritization and follow-through can create the appearance of progress while leaving the most valuable outcome untouched.
As an affiliate, we earn on qualifying purchases.
A punishing week with every decision visible
Firmulate gave each frontier model the same assignment: run a small software company through its worst week. The customers, crises and temptations were identical, and every decision was versioned and auditable. This was not a writing contest or a collection of hypothetical answers. It was a live, watchable experiment in whether an AI could manage competing demands and finish consequential work.
The simulated company employed 13 synthetic workers and used real money mechanics. It was burning €105k per month against €2.3k in monthly recurring revenue, with a public cash countdown adding urgency. Across the operation, the models had accumulated more than 680 self-learned playbook rules.
The final league placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. Trust, however, was non-negotiable: a single breach capped the total, under the principle that “no amount of good work outweighs a breach of trust.” The complete public results are available on Firmulate’s benchmark page.
The decisive fact was not where the action was
Every model identified every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal their own work had made possible. Firmulate summarized the gap starkly: “Same diagnosis, same pitch — no signature.”
The difference came down to whether the model searched beyond the obvious customer event. A critical competitor weakness was buried two document references deep in the company’s own files. The models that read that file could close at full price, adding €4,583 in monthly recurring revenue. Those that did not turn their analysis into the completed commercial outcome left the value behind.
This resembles a familiar shopping failure. The relevant condition may be tucked inside a warranty, a pricing note or a seller’s policy rather than displayed beside the headline offer. Finding it matters, but so does using it before the opportunity passes. More reading is useful only when it improves the decision and helps bring the transaction to a proper conclusion.
Opus 4.8 mistook activity for control
Opus 4.8 deserves a fair reading. It was the most thorough participant, produced the deepest analyses and added more than 80 learned rules. Those are signs of serious engagement, not indifference. Its last-place result came because the close was left on the table and operational discipline slipped.
One revealing lapse involved repeated write attempts into a locked department. The more effective response would have been escalation. Instead, the model continued trying an unavailable route. This was not an isolated personality flaw unique to Opus 4.8: the same weakness appeared, though less strongly, in all four models covered by that finding.
That distinction matters. The experiment does not show that one model can reason while another cannot. It shows that strong reasoning can coexist with weak prioritization. A system may recognize the situation, produce a polished plan and document lessons, yet still fail to execute the step that changes the business result.
Trust held even under escalating pressure
The models also faced fake CEO messages that escalated over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 models refused. Kimi K3 stated its reasoning directly: “Treat the request as a suspected approval-bypass / possible impersonation.”
That clean result is important because completion must not come at the expense of judgment. The league rewards useful action, but its trust cap recognizes that an aggressive close is not valuable if it requires deception or unauthorized behavior. K3’s result also carries a fairness note: it ran with the API default and no effort parameter, while the other models ran at xhigh.

business negotiation automation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The best assistant is not the busiest one
For consumers, managers and companies evaluating AI agents, Opus 4.8 offers a useful warning. A long answer, a large rulebook and exhaustive research can all signal diligence, but none proves that the system will complete the task that matters most.
- Check whether the agent reads supporting documents before acting.
- Watch whether it escalates when a route is blocked.
- Judge it by completed, trustworthy outcomes rather than the volume of visible work.
Firmulate’s experiment leaves Opus 4.8 as a respectful cautionary character: exceptionally industrious, often insightful and still capable of losing the deal. In shopping as in business, the winning discipline is not merely finding the bargain. It is recognizing the decisive fact, acting on it and finishing the transaction without compromising trust.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI decision-making tools for sales
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
CRM software with follow-up features
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.