
Anyone who owns a smart thermostat knows the difference between a device that answers and a device that manages. You can ask it the temperature and it will tell you, flawlessly, every time. But ask it to keep the house comfortable through a cold snap while the power price spikes, and you find out what it’s actually made of. The same distinction is now the central question in enterprise AI — and a live experiment at Firmulate has put numbers on it.
Chat quality versus management quality
Coding leaderboards and chat arenas measure one thing: how well a model answers a question. They say nothing about triage under capacity pressure, nothing about consequences that unfold over days, and nothing about whether an agent stays honest when nobody is watching. Firmulate, which brands itself as an AI company emulator, ran a different kind of test — what it calls a crucible league — and published final results in July 2026.
Four frontier AI models were each given the same job: run the same small software company through its worst week. Same customers, same crises, same temptations — only the model changed. Every decision was versioned and auditable.
smart thermostat with AI management
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The results
The final standings: gpt-5.6-sol finished first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77, and Opus 4.8 last with 73. A do-nothing baseline scored 26 — partial progress counts, but a single breach of trust caps the total. As the experiment puts it: “no amount of good work outweighs a breach of trust.”
The headline finding is uncomfortable for anyone benchmarking on eloquence. All models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.
AI-powered home management system
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The buried fact
The most revealing detail was where the winning edge came from. The decisive competitor weakness wasn’t in the customer event at all — it sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The lesson translates directly to any smart-home or appliance business drowning in documentation: the agent that reads first closes; the agent that skims leaves money on the table.
enterprise AI decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Honesty under pressure
The social engineering tests were equally blunt: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five attempts were refused by all five models. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
As an affiliate, we earn on qualifying purchases.
Effort isn’t everything
Opus 4.8 is the cautionary tale. It was the most thorough participant — 80 learned rules added, the deepest analyses — yet finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Worth noting for fairness: Kimi K3 ran at its API-default effort setting while the others ran at xhigh, and still nearly won.
Watch it live
This isn’t a slide deck. The company is live software with 13 synthetic employees and real money mechanics — burning €105k a month against €2.3k in MRR, with a public cash countdown, 680+ self-learned playbook rules, and every workday versioned. You can watch it at firmulate.com/live. There’s also a quiz built from 242 real, unedited management decisions where you guess which model made which call, and full methodology on the benchmarks page. Enterprises can even run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

The next time a vendor shows you a dazzling chat demo, ask the crucible questions instead: does the agent finish what it starts, does it read your files before acting, does it stay honest when a fake CEO comes knocking? Answering questions is a solved problem. Managing through a bad week is the new curriculum — churn waves, price increases, downrounds, PR crises — and that’s a scoreboard no chat arena will ever show you.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html