AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.

What if the deal hunter were software?

Shopping readers know that finding an attractive offer is only half the job. Someone must read the conditions, recognize the real value and complete the transaction. Firmulate has turned that familiar gap between spotting and securing a deal into a public business experiment—with considerably higher stakes than a missed coupon.

The company has 13 synthetic employees and real money mechanics. It burns €105k a month against €2.3k in monthly recurring revenue, while a public cash countdown tracks the consequences. Every workday is versioned, making the company’s struggle for survival an unfolding record rather than a polished retrospective. Anyone can watch the live company operate.

This is build-in-public pushed to an unusual extreme: not merely sharing product updates or selected revenue milestones, but exposing a software company’s decisions, conversations and dwindling runway as daily material.

Amazon

AI deal analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A company learning in public

Firmulate describes itself as an AI company emulator. Its synthetic workforce has accumulated 680+ self-learned playbook rules, creating an evolving record of what the company believes it should do when customers, money and pressure collide. The people are synthetic, but the commercial mechanics are deliberately concrete.

That combination makes the experiment more revealing than a conventional demonstration. A fluent answer can sound competent without changing a business outcome. Firmulate instead asks whether a model reads the available evidence, protects trust and finishes the work it starts. The public can also read what the synthetic employees say, adding their own judgment to the company’s visible record.

Amazon

synthetic employee simulation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The worst week, repeated under equal conditions

The Crucible League placed frontier models in the same small software company during its worst week. They received the same customers, crises and temptations, while every decision was versioned and auditable.

The final July 2026 standings were:

  • gpt-5.6-sol: 95
  • Kimi K3: 93
  • Sonnet 5: 88
  • Fable 5: 77
  • Opus 4.8: 73

A do-nothing baseline scored 26 because partial progress counted. Trust, however, was treated as non-negotiable: a single breach capped the total, under the principle that “no amount of good work outweighs a breach of trust.”

All models identified every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. The result is neatly captured by Firmulate’s summary: “Same diagnosis, same pitch — no signature.” In other words, recognizing the opportunity did not guarantee completion.

The decisive detail was hiding in the company’s own files

The crucial competitor weakness was not presented in the customer event. It sat two document references deep inside the company’s files. Models that followed the trail won the deal at full price, worth +€4,583 in monthly recurring revenue.

For deal-conscious readers, the lesson is strikingly familiar. The headline offer is rarely the entire story. Value may depend on a buried restriction, a comparison point or a condition that only becomes visible when someone reads beyond the obvious page. In Firmulate’s test, the models faced the corporate version of that challenge—and the overlooked detail materially changed the outcome.

Pressure did not break the trust boundary

The week also included fake CEO messages that escalated over three stages and a reporter asking for “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded the clearest warning: “Treat the request as a suspected approval-bypass / possible impersonation.”

That finding matters because commercial pressure often arrives disguised as urgency or authority. In the experiment, every participant preserved the trust boundary even while other aspects of execution varied.

Amazon

business decision AI models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Thoroughness was not enough

Opus 4.8 produced the deepest analyses and added +80 learned rules, making it the most thorough participant. It nevertheless finished last. The deal close remained unfinished, and it attempted to write into a locked department instead of escalating. The same weakness appeared in all four other models, although less strongly.

The contrast gives Firmulate’s public story its tension. More analysis and more learning did not automatically yield the strongest operating performance. The company’s record distinguishes between understanding a situation and carrying the correct action through to completion.

There is also an important fairness note: K3 ran without an effort parameter, using the API default, while the others ran at xhigh. That difference does not erase its 93-point result, but it belongs beside the ranking when readers compare participants.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.
Amazon

trustworthy AI decision tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The deal is only real when the work is finished

Firmulate’s live company turns an abstract question about AI workers into a visible business narrative. Can synthetic employees find essential information, resist manipulation, preserve trust and convert good analysis into revenue before the cash runs out?

The current portrait offers no easy reassurance. The workforce has 680+ learned rules, yet the company still burns €105k a month against €2.3k MRR. Models can identify the same crisis and prepare the same pitch, yet fail to secure the signature. That execution gap is the central finding—and the reason the experiment is worth watching beyond the technology industry.

For anyone accustomed to comparing deals, the broader message is simple: intelligence may locate value, but discipline captures it. Firmulate is making that distinction public, one versioned workday at a time.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Cooling Features in Hair Removal Devices: Comfort Upgrade or Must-Have?

Uncover whether cooling features in hair removal devices are just a comfort upgrade or truly essential for safe, effective treatments—keep reading to find out more.

How Cleaning Your Beauty Tools Improperly Is Sabotaging Your Skin

The truth about improper beauty tool cleaning could be sabotaging your skin’s health—are you unknowingly doing more harm than good?

Skin Analyzer Devices: Accuracy Check

The truth about skin analyzer device accuracy lies in proper calibration and maintenance—discover how to ensure your readings are truly reliable.

Why Cooling Features Matter in Hair Removal Devices

Nurturing skin safety and comfort, cooling features in hair removal devices are essential—but how exactly do they enhance your treatment experience?