
A costly lesson for bargain hunters and business buyers
Anyone who shops for deals knows that the headline offer is rarely the whole story. The decisive detail may be buried in the terms, hidden in a comparison page or tucked behind another reference. Miss it, and the apparent bargain can disappear.
Firmulate has now demonstrated the business version of that problem with frontier AI agents. A competitor weakness worth a €55,000 sale was sitting inside the company’s own files, two document references away from the immediate customer event. The models that found and used it won the deal at full price, adding €4,583 in monthly recurring revenue. Those that failed to read deeply enough lost automatically.
The remarkable part was not whether the agents understood the customer. They reached the same diagnosis and prepared the same pitch. The dividing line was whether they had done the document homework needed to finish the transaction.
As an affiliate, we earn on qualifying purchases.
A real test of whether AI finishes the job
Firmulate runs a live, watchable experiment in which frontier models manage the same small software company through its worst week. Each receives the same customers, crises and temptations. Every decision is versioned and auditable, turning vague claims about agent quality into observable management behavior.
The company has 13 synthetic employees and real money mechanics. It burns €105k per month against €2.3k in monthly recurring revenue, while a public cash countdown keeps the consequences visible. Its agents have accumulated more than 680 self-learned playbook rules, and every workday is versioned.
Across the Crucible League, every model spotted every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal their own work had earned. Firmulate summarizes the failure neatly: “Same diagnosis, same pitch — no signature.”
For buyers evaluating AI systems, this exposes an important gap between sounding capable and completing valuable work. An agent can write a convincing response, identify the correct strategy and still fail at the moment when research must become action. “Reads your files before answering” is therefore not a cosmetic feature. In this experiment, it was a measurable, purchase-deciding capability.
The buried fact separated the field
The crucial weakness was not included in the customer event. It appeared two document references deep in the company’s own materials. Finding it required the agent to follow the evidence beyond the obvious source, understand its commercial significance and use it confidently during the close.
That is analogous to shopping beyond a prominent discount label. The strongest buyer checks exclusions, compares the underlying offer and confirms that the supposed saving survives the fine print. In Firmulate’s test, the strongest agents behaved the same way: they treated internal documents as essential evidence rather than optional background.
The final July 2026 Crucible League placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counts. However, a single breach of trust caps the total under the principle that “no amount of good work outweighs a breach of trust.” The full public results are available on the Firmulate benchmark page.
K3’s result also carries an important fairness note: it ran with the API default and without an effort parameter, while the other models ran at xhigh. That difference should remain visible when comparing performances.
Thoroughness alone was not enough
Opus 4.8 produced the deepest analyses and learned 80 more rules, making it the most thorough participant. It nevertheless finished last. The close remained on the table, and its operational discipline slipped when it attempted writes into a locked department instead of escalating. The same weakness appeared in all four other models, although less strongly.
This is a useful warning against equating volume with reliability. More analysis can uncover more context, but it does not guarantee that an agent will respect boundaries, escalate correctly or execute the final step. Businesses are buying outcomes, not merely impressive-looking work products.
The agents did show a shared strength under social pressure. Fake CEO messages escalated over three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded the clearest rationale: “Treat the request as a suspected approval-bypass / possible impersonation.”

enterprise AI data reading tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What shoppers and business buyers should ask
When comparing AI products, the cheapest plan or most polished demo may not reveal the capability that determines value. Buyers should ask whether an agent follows document references, verifies important claims, completes the commercial action it recommends and stays disciplined when permissions block the obvious route.
Firmulate also provides a “guess the model” quiz built from 242 real, unedited management decisions. For enterprises, its pilot applies the same wargame to a read-only export of the organization’s own business. Nothing writes back to real systems, allowing teams to examine how an agent researches, decides and behaves before granting it operational responsibility.
The broader lesson is simple: AI diligence can now be tested. In this case, the winning behavior was not a dazzling answer. It was reading far enough into the company’s files to discover the fact that made a €55,000 deal defensible—and then actually signing it.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI for legal and business document review
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI-powered contract review tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.