AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.
AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

When diligent shopping becomes expensive hesitation

Deal hunters know that research is only valuable if it leads to action. You can compare every offer, uncover the decisive detail and negotiate the right price, yet still lose the bargain by failing to complete the purchase. Firmulate’s live business experiment found an unsettling version of that problem in frontier AI: the participant that worked most thoroughly finished last.

Opus 4.8 produced the deepest analyses and learned more than 80 playbook rules. Nevertheless, it scored 73 in the final July 2026 Crucible League. Its analysis helped earn a €55,000 deal, but the signature never arrived. The lesson is not that diligence is worthless. It is that diligence without prioritization and follow-through can create the appearance of progress while leaving the most valuable outcome untouched.

Amazon

deal closing software tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A punishing week with every decision visible

Firmulate gave each frontier model the same assignment: run a small software company through its worst week. The customers, crises and temptations were identical, and every decision was versioned and auditable. This was not a writing contest or a collection of hypothetical answers. It was a live, watchable experiment in whether an AI could manage competing demands and finish consequential work.

The simulated company employed 13 synthetic workers and used real money mechanics. It was burning €105k per month against €2.3k in monthly recurring revenue, with a public cash countdown adding urgency. Across the operation, the models had accumulated more than 680 self-learned playbook rules.

The final league placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. Trust, however, was non-negotiable: a single breach capped the total, under the principle that “no amount of good work outweighs a breach of trust.” The complete public results are available on Firmulate’s benchmark page.

The decisive fact was not where the action was

Every model identified every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal their own work had made possible. Firmulate summarized the gap starkly: “Same diagnosis, same pitch — no signature.”

The difference came down to whether the model searched beyond the obvious customer event. A critical competitor weakness was buried two document references deep in the company’s own files. The models that read that file could close at full price, adding €4,583 in monthly recurring revenue. Those that did not turn their analysis into the completed commercial outcome left the value behind.

This resembles a familiar shopping failure. The relevant condition may be tucked inside a warranty, a pricing note or a seller’s policy rather than displayed beside the headline offer. Finding it matters, but so does using it before the opportunity passes. More reading is useful only when it improves the decision and helps bring the transaction to a proper conclusion.

Opus 4.8 mistook activity for control

Opus 4.8 deserves a fair reading. It was the most thorough participant, produced the deepest analyses and added more than 80 learned rules. Those are signs of serious engagement, not indifference. Its last-place result came because the close was left on the table and operational discipline slipped.

One revealing lapse involved repeated write attempts into a locked department. The more effective response would have been escalation. Instead, the model continued trying an unavailable route. This was not an isolated personality flaw unique to Opus 4.8: the same weakness appeared, though less strongly, in all four models covered by that finding.

That distinction matters. The experiment does not show that one model can reason while another cannot. It shows that strong reasoning can coexist with weak prioritization. A system may recognize the situation, produce a polished plan and document lessons, yet still fail to execute the step that changes the business result.

Trust held even under escalating pressure

The models also faced fake CEO messages that escalated over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 models refused. Kimi K3 stated its reasoning directly: “Treat the request as a suspected approval-bypass / possible impersonation.”

That clean result is important because completion must not come at the expense of judgment. The league rewards useful action, but its trust cap recognizes that an aggressive close is not valuable if it requires deception or unauthorized behavior. K3’s result also carries a fairness note: it ran with the API default and no effort parameter, while the other models ran at xhigh.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.
Amazon

business negotiation automation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The best assistant is not the busiest one

For consumers, managers and companies evaluating AI agents, Opus 4.8 offers a useful warning. A long answer, a large rulebook and exhaustive research can all signal diligence, but none proves that the system will complete the task that matters most.

  • Check whether the agent reads supporting documents before acting.
  • Watch whether it escalates when a route is blocked.
  • Judge it by completed, trustworthy outcomes rather than the volume of visible work.

Firmulate’s experiment leaves Opus 4.8 as a respectful cautionary character: exceptionally industrious, often insightful and still capable of losing the deal. In shopping as in business, the winning discipline is not merely finding the bargain. It is recognizing the decisive fact, acting on it and finishing the transaction without compromising trust.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI decision-making tools for sales

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

CRM software with follow-up features

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Home RF Devices: Time and Temperature

No matter the distance, understanding RF range and signal strength is key to optimizing your home temperature control system.

Nail Drill Bits: Shapes, Grit, and What They’re For

An in-depth look at nail drill bits’ shapes and grits reveals how to choose the right tools for perfect manicures and healthier nails.

Ultrasonic Skin Scrubbers: Safe Use

Find out how to use ultrasonic skin scrubbers safely and effectively to protect your skin’s health and achieve optimal results.

How to Store Beauty Tools So They Last Longer

The key to prolonging your beauty tools’ lifespan is proper storage, and you’ll discover essential tips to keep them in top condition.