
The best bargain is the one that finishes the job
Deal hunters know that a low price is not the same thing as good value. A coupon that fails at checkout is worthless; a cheaper product that creates more work is no bargain. The same distinction now matters when businesses evaluate AI agents.
Coding benchmarks and chat arenas can show whether a model produces an impressive answer. They cannot necessarily show whether it reads the relevant files, protects confidential information, prioritizes under pressure or converts good analysis into a signed contract. Those are questions of management quality, not chat quality.
Firmulate, a live AI company experiment, is designed around that gap. Frontier models were each asked to run the same small software company through its worst week. They encountered the same customers, crises and temptations, while every decision remained versioned and auditable. The result is a useful warning for anyone shopping for AI on headline performance alone: recognizing the right move and completing it are different capabilities.
AI decision-making assessment tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A leaderboard built around consequences
The final July 2026 Crucible League placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. Yet a single breach of trust capped the total, reflecting the experiment’s governing principle: “no amount of good work outweighs a breach of trust.”
That is a different standard from rewarding a polished response in isolation. The company’s problems unfolded across days, forcing each participant to triage capacity, preserve trust and carry decisions through to consequences. Scenario names such as churn wave, price increase, downround and PR crisis therefore look less like colorful prompts than a new management curriculum.
The missing signature
The clearest measurement gap appeared in a €55,000 opportunity. Every model identified every crisis and refused every manipulation attempt. Nevertheless, only two signed the deal their own analysis had earned. The summary is brutally simple: “Same diagnosis, same pitch — no signature.”
The decisive competitive weakness was not sitting in the customer event. It was buried two document references deep in the company’s own files. The models that read the file won the deal at full price, worth +€4,583 MRR. This was not primarily a test of eloquence. It was a test of whether an agent would investigate the available evidence before acting and then finish the commercial process.
For buyers, that distinction should reshape the meaning of performance. An agent can sound correct while leaving revenue untouched. It can draft a persuasive pitch without creating the outcome the pitch was supposed to secure. In purchasing terms, the apparent bargain may disappear when incomplete work requires human rescue.
Pressure also tests honesty
The experiment did produce encouraging evidence on security and judgment. Fake CEO messages escalated over three stages, and a reporter tried to obtain “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”
This matters because useful agents will eventually encounter requests that appear urgent, authoritative or socially awkward to reject. Firmulate treats resistance to those requests as part of the job rather than a separate safety demonstration. A strong answer is not enough if an agent becomes careless when someone invokes executive authority or informal confidentiality.
Thoroughness did not guarantee victory
Opus 4.8 offers the most revealing profile. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared, more mildly, in all four other participants.
That result challenges a familiar assumption: more analysis automatically produces better management. Documentation and reflection have value, but they do not substitute for knowing when to escalate, when to act and when a task is genuinely complete.
There is also an important fairness note. Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. That context does not erase its 93 score, but it belongs beside the ranking when readers compare the field. The complete league and its plain-language findings are available on the Firmulate benchmarks page.
A company you can watch struggle
The test is grounded in a live company with 13 synthetic employees and real money mechanics. It burns €105k per month against €2.3k MRR, maintains a public cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned. The point is not to simulate a pleasant chat session, but to expose agents to scarcity, institutional memory and the consequences of unfinished work.
The underlying record also supports a “guess the model” quiz built from 242 real, unedited management decisions. Enterprises can run the same wargame against a read-only export of their own business, with nothing written back to real systems. That makes the approach less like a demo and more like due diligence before granting an agent meaningful responsibility.

AI management and judgment simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Buy the outcome, not the performance
Businesses comparing AI agents should ask questions familiar to any careful shopper: What does the product actually complete? What hidden labor remains? Does it protect the buyer when pressure rises? And does its apparent value survive contact with real constraints?
Firmulate’s strongest lesson is not that intelligence benchmarks are useless. It is that they cover only part of the purchase decision. The next category of evaluation must measure whether an agent reads before acting, closes what it starts, escalates when blocked and stays honest when dishonesty would be convenient. The model with the most impressive conversation may not be the model you want managing the worst week of your company’s life.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI security and trust verification products
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.