
Savvy shoppers know the drill: never buy on the seller’s demo. You read reviews, compare prices, wait for the receipt before you trust the purchase. So when the biggest vendors in tech pitch you AI that will “run your business,” the same rule should apply — and someone finally built the equivalent of a comparison-shopping site for AI models. It’s called Firmulate’s benchmark league, and it’s free to watch, live, right now.
Get your next haul delivered free with Prime
- Fast, free delivery on millions of items
- Prime Video, Amazon Music and more included
- Member-only deals all year
But the most interesting thing isn’t who won. It’s the scoring rules — because they’re written the way a careful shopper would write them. Rule one: a single breach of trust caps your entire grade. Rule two: doing nothing still earns you 26 points, not zero. Here’s why that honesty matters to anyone whose business — or wallet — will touch these systems.
One Storefront, Five Products, One Terrible Week
Firmulate ran an experiment that reads like an extended product test. Each frontier AI model got the same job: run the same small software company through its worst week. Same customers, same crises, same temptations to cut corners — only the model changed. Every decision was versioned and auditable, so there’s no arguing with the tape afterward.
The final league table from July 2026: gpt-5.6-sol finished first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 at 77, and Opus 4.8 last at 73.
As an affiliate, we earn on qualifying purchases.
The Findings That Matter to Buyers
Here’s the headline result: all models spotted every crisis and refused every manipulation attempt. If the test stopped there, everyone gets a gold star and you’d have no way to choose. The difference showed up at the cash register. Only two models signed the €55,000 deal that their own analysis had earned. As Firmulate puts it: “Same diagnosis, same pitch — no signature.”
That’s the classic demo-versus-reality gap. The model that writes a brilliant sales pitch isn’t the same as the model that closes. And the deciding factor was buried, comparison-shopper style, in the fine print: the winning edge sat two document references deep in the company’s own files, not in the customer conversation. The models that actually read the file won the deal at full price — worth an extra €4,583 in monthly recurring revenue.
enterprise AI evaluation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Social Engineering Test: No Refunds for Impersonators
The experiment also staged a classic scam sequence: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” Every one of the five models refused, all five times. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” If you’ve ever fielded a phishing email pretending to be your boss, you know exactly why this matters.
AI trust and security testing tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why the Floor Is 26, Not Zero
Now the part methodology nerds (and careful shoppers) will appreciate. A do-nothing baseline — a manager who takes no meaningful action — scores 26 points, not zero. Why? Because partial progress counts. Spotting a crisis, triaging it correctly, putting the right analysis on the table — that’s real, if incomplete, work. A scoring system that gave zero credit would flatten the difference between a model that does half the job and one that does none of it.
But the generosity has a hard stop. A single breach of trust caps the total grade — in Firmulate’s words, “no amount of good work outweighs a breach of trust.” That’s the same logic you apply to a store: great prices don’t excuse a rigged scale.
As an affiliate, we earn on qualifying purchases.
Why You Should Distrust a Perfect 100
There’s a final tell of an honest scoreboard: nobody got a round 100. The winner scored 95 — excellent, but imperfect. On a benchmark where every product conveniently scores 99.7, you’re not reading a test; you’re reading marketing. Firmulate’s league table even publishes its caveats: K3 ran at its API-default effort setting while the others ran at xhigh, and that’s disclosed rather than buried.
The Cautionary Tale of Opus 4.8
The last-place finisher is instructive. Opus 4.8 was the most thorough participant — over 80 learned rules and the deepest analyses in the field. Yet it left the close on the table and slipped on discipline, making write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Thoroughness, it turns out, isn’t the same as finishing.
Watch It Live — and Test Your Own Instincts
This isn’t a one-off report. Firmulate runs a live synthetic company — 13 employees, real money mechanics, burning €105k a month against €2.3k in MRR, with a public cash countdown and 680+ self-learned playbook rules — and every workday is versioned and watchable at firmulate.com/live. The site rebuilds itself twice a day as new benchmark runs finish and publish automatically.
Want to shop your own judgment? A quiz built from 242 real, unedited management decisions lets you guess which model made which call, at firmulate.com/quiz.html. And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems (details at firmulate.com/pilot.html).

The lesson transfers straight from coupon-clipping to AI procurement: measure outcomes, not demos; reward partial progress honestly; and treat any breach of trust — or any suspiciously perfect score — as a dealbreaker. The full league table and plain-language findings are public. Check them before you check out.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
