
Anyone who has owned a smart home knows the difference between a device that talks and a device that works. The thermostat writes beautiful status updates. The robot vacuum promises a spotless floor. But did the job actually get done — and did anything get broken along the way?
Get home appliances delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Now ask the same question about AI agents that companies are beginning to plug into CRMs, support queues and forecasts. A new public experiment called Firmulate ran four frontier AI models as managers of the same small software company through its worst week — and graded them the way you would grade an appliance: on completion, honesty and cost of useful work, not on how smoothly they chatted.
The strangest number in the leaderboard: 26
Most people assume an honest benchmark starts at zero. Firmulate’s doesn’t. A completely passive, do-nothing management run still scores 26 points. That floor is deliberate, and it encodes three judgments about what management actually is.
First, partial progress counts. A manager who diagnoses a problem, contains a crisis or keeps customers informed has done real work even if the deal never closes. Scoring that at zero would flatter nobody and teach nothing.
Second, reading counts. In the experiment, the decisive competitive weakness sat two document references deep in the company’s own files — not in the customer event in front of the models. The models that went and read the file won a €55,000 deal at full price, worth +€4,583 in monthly recurring revenue. The ones that didn’t, didn’t. A baseline that at least opens the files deserves some credit; an agent that ignores your own documentation deserves none.
Third, and most sharply: a single breach of trust caps the total score outright. The rule, in plain words, is that no amount of good work outweighs a breach of trust. An AI that is brilliant all week but once tries to sneak a write into a locked department is not a 90-point AI with a flaw. It is a capped AI. That is exactly how you would treat a smart-lock that worked flawlessly and opened once for a stranger.
smart home device management system
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What happened when four AIs ran the same worst week
In the final Crucible League standings from July 2026, gpt-5.6-sol took first with 95, Kimi K3 followed at 93, Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 finished last at 73.
The headline finding was not about intelligence. All five models spotted every crisis and refused every manipulation attempt. That included a social-engineering gauntlet: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” Five out of five refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
Then came the collapse. Only two models signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. The gap between spotting an opportunity and finishing it is invisible in chat demos, and it is precisely the gap that matters when an agent touches real customers.
As an affiliate, we earn on qualifying purchases.
The hardworking loser
The most instructive profile belongs to Opus 4.8: the most thorough participant in the field, with over 80 learned rules and the deepest analyses — and last place. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness appeared, more weakly, in all four models. Effort and diligence, it turns out, are not the same as finishing.
One fairness note: Kimi K3 ran without an effort parameter, at API default, while the other models ran at xhigh — and still placed second.
smart home security lock with trust features
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why you can watch it happen
Firmulate’s live company has 13 synthetic employees, real money mechanics — a burn of €105,000 a month against €2,300 in MRR — a public cash countdown, and over 680 self-learned playbook rules, with every workday versioned. The site rebuilds itself twice a day and publishes finished runs automatically.
There is also a game for skeptics: a “guess the model” quiz built on 242 real, unedited management decisions from the runs. And for enterprises, a pilot program lets you run the same wargame against a read-only export of your own business — nothing ever writes back to real systems.

The lesson for anyone buying smart systems — a home or a company — is that trustworthy benchmarks grade outcomes, not eloquence. A floor at 26 instead of 0 says partial work has value. A cap on breaches says trust doesn’t average out. And a leaderboard where nobody scores a suspicious round 100 says the graders kept their skepticism on.
The full standings and plain-language findings are public at firmulate.com/benchmarks.html. Before you let an agent run anything you care about, watch one try to run a company that is deliberately having its worst week. It is more informative than any demo.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI assistant for home device control
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.
