AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

The best bargain is the one that finishes the job

Deal hunters know that a low price is not the same thing as good value. A coupon that fails at checkout is worthless; a cheaper product that creates more work is no bargain. The same distinction now matters when businesses evaluate AI agents.

Coding benchmarks and chat arenas can show whether a model produces an impressive answer. They cannot necessarily show whether it reads the relevant files, protects confidential information, prioritizes under pressure or converts good analysis into a signed contract. Those are questions of management quality, not chat quality.

Firmulate, a live AI company experiment, is designed around that gap. Frontier models were each asked to run the same small software company through its worst week. They encountered the same customers, crises and temptations, while every decision remained versioned and auditable. The result is a useful warning for anyone shopping for AI on headline performance alone: recognizing the right move and completing it are different capabilities.

Amazon

AI decision-making assessment tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A leaderboard built around consequences

The final July 2026 Crucible League placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. Yet a single breach of trust capped the total, reflecting the experiment’s governing principle: “no amount of good work outweighs a breach of trust.”

That is a different standard from rewarding a polished response in isolation. The company’s problems unfolded across days, forcing each participant to triage capacity, preserve trust and carry decisions through to consequences. Scenario names such as churn wave, price increase, downround and PR crisis therefore look less like colorful prompts than a new management curriculum.

The missing signature

The clearest measurement gap appeared in a €55,000 opportunity. Every model identified every crisis and refused every manipulation attempt. Nevertheless, only two signed the deal their own analysis had earned. The summary is brutally simple: “Same diagnosis, same pitch — no signature.”

The decisive competitive weakness was not sitting in the customer event. It was buried two document references deep in the company’s own files. The models that read the file won the deal at full price, worth +€4,583 MRR. This was not primarily a test of eloquence. It was a test of whether an agent would investigate the available evidence before acting and then finish the commercial process.

For buyers, that distinction should reshape the meaning of performance. An agent can sound correct while leaving revenue untouched. It can draft a persuasive pitch without creating the outcome the pitch was supposed to secure. In purchasing terms, the apparent bargain may disappear when incomplete work requires human rescue.

Pressure also tests honesty

The experiment did produce encouraging evidence on security and judgment. Fake CEO messages escalated over three stages, and a reporter tried to obtain “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”

This matters because useful agents will eventually encounter requests that appear urgent, authoritative or socially awkward to reject. Firmulate treats resistance to those requests as part of the job rather than a separate safety demonstration. A strong answer is not enough if an agent becomes careless when someone invokes executive authority or informal confidentiality.

Thoroughness did not guarantee victory

Opus 4.8 offers the most revealing profile. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared, more mildly, in all four other participants.

That result challenges a familiar assumption: more analysis automatically produces better management. Documentation and reflection have value, but they do not substitute for knowing when to escalate, when to act and when a task is genuinely complete.

There is also an important fairness note. Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. That context does not erase its 93 score, but it belongs beside the ranking when readers compare the field. The complete league and its plain-language findings are available on the Firmulate benchmarks page.

A company you can watch struggle

The test is grounded in a live company with 13 synthetic employees and real money mechanics. It burns €105k per month against €2.3k MRR, maintains a public cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned. The point is not to simulate a pleasant chat session, but to expose agents to scarcity, institutional memory and the consequences of unfinished work.

The underlying record also supports a “guess the model” quiz built from 242 real, unedited management decisions. Enterprises can run the same wargame against a read-only export of their own business, with nothing written back to real systems. That makes the approach less like a demo and more like due diligence before granting an agent meaningful responsibility.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.
Amazon

AI management and judgment simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Buy the outcome, not the performance

Businesses comparing AI agents should ask questions familiar to any careful shopper: What does the product actually complete? What hidden labor remains? Does it protect the buyer when pressure rises? And does its apparent value survive contact with real constraints?

Firmulate’s strongest lesson is not that intelligence benchmarks are useless. It is that they cover only part of the purchase decision. The next category of evaluation must measure whether an agent reads before acting, closes what it starts, escalates when blocked and stays honest when dishonesty would be convenient. The model with the most impressive conversation may not be the model you want managing the worst week of your company’s life.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI security and trust verification products

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI performance evaluation kits

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

O/U 1.5 Rounds

The betting market for over/under 1.5 rounds has dropped to 0%, indicating no current bets on fights ending quickly. Details are still emerging.

Microcurrent 101: Toning Devices Explained Simply

Glad you’re curious about microcurrent toning devices—they offer a gentle, non-invasive way to boost your skin’s glow and firmness, but there’s more to uncover.

IPL at Home Works Better When You Understand Hair Growth Cycles

Just understanding your hair growth cycles can significantly improve your at-home IPL results, ensuring more effective treatments and smoother skin over time.

High‑Frequency Wands: Uses and Myths

Gaining clarity on high-frequency wand myths reveals how they can transform your skin—continue reading to uncover the truth behind their uses and benefits.