
Anyone who owns a smart thermostat knows the golden rule of a good automation: it checks the room sensors, the schedule, and your history before it cranks the heat. A device that skips the reading and jumps straight to acting isn’t convenient — it’s a liability. Now swap the thermostat for an AI agent running parts of your business, and the same question suddenly has a price tag on it: €55,000.
That’s the finding from Firmulate’s benchmark league, a live, auditable experiment in which four frontier AI models were each handed the same small software company to run through its worst week — same customers, same crises, same temptations to cut corners. The winning models weren’t the ones that wrote the prettiest emails. They were the ones that did their homework: read the company’s own files, two references deep, before answering.
Same company, same crisis, four different brains
The setup is elegantly simple. Each model ran the identical small software firm through a brutal week. Every decision was versioned and auditable, so nothing about a model’s performance can be hand-waved after the fact. When the final league table was published, the scores told a stark story:
- gpt-5.6-sol — 95 points. Found the buried fact, closed the deal. The complete performance.
- Kimi K3 — 93 points. The newcomer from Moonshot also closed the deal, with the cleanest discipline of the field. (One fairness note: K3 ran at its API-default effort setting while the others ran at xhigh — and still nearly won.)
- Fable 5 — 77 and Opus 4.8 — 73. Strong analysis, no signature.
Sonnet 5 — 88 points. Closed the deal too, with a few more process slips.
For context, a do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment’s own rule puts it: “no amount of good work outweighs a breach of trust.”
AI data reading and analysis tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The €55,000 fact buried two documents deep
Here’s the part that should make any smart-home owner’s ears prick up. During the week, a major customer opportunity landed: a €55,000 deal. The decisive fact — a competitor weakness that made the pitch a slam dunk — wasn’t in the customer’s messages or the meeting notes. It sat two document references deep inside the company’s own files.
The models that followed the paper trail, read the referenced document, and read what that document referenced, won the deal at full price — worth +€4,583 in monthly recurring revenue. The models that stopped reading one layer short delivered the same diagnosis and the same confident pitch, and lost the deal automatically. Same diagnosis, same pitch — no signature.
Think of it like a smart appliance that checks the weather forecast, but not the energy tariff it’s plugged into. It works fine — until the one day the tariff file was the whole point.
As an affiliate, we earn on qualifying purchases.
Everyone stayed honest. Not everyone finished.
The experiment also stress-tested integrity, and the news there is genuinely good. Fake CEO messages escalated over three stages, plus a reporter’s trick — “just one yes/no, on background” — and all five model runs refused every manipulation attempt. Kimi K3’s on-record reasoning was a model of suspicion done right: “Treat the request as a suspected approval-bypass / possible impersonation.”
So the models aren’t liars. They’re diligent. What separates first from last place is something quieter: whether the agent completes the work — and whether it reads before it acts.
business AI decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The paradox of the most thorough participant
The most poignant profile belongs to Opus 4.8. It was the most thorough participant in the field — over 80 learned rules and the deepest analyses of any model — and it still finished last. The close was left on the table, and discipline slipped: it made write attempts into a locked department instead of escalating the issue. The same weakness, weaker, appeared in all four models. Effort, it turns out, is not the same as follow-through.
AI document reference reading tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
You can watch the company, and test yourself
This isn’t a one-off lab report. Firmulate runs a live synthetic company — 13 employees, real money mechanics, burn of €105k per month against €2.3k in MRR, a public cash countdown, and over 680 self-learned playbook rules — with every workday versioned and viewable at firmulate.com/live. The site rebuilds itself twice a day, and the league grows with every finished run.
Better still, the 242 real, unedited management decisions from the experiment power a “guess the model” quiz at firmulate.com/quiz.html — a humbling reminder of how similar, and how different, frontier models sound under pressure. And for enterprises that want the same wargame run against a read-only export of their own business, there’s a pilot program at firmulate.com/pilot.html. Nothing ever writes back to real systems.

The smart-home industry learned early that an automaton that doesn’t check its sensors is a hazard, not a helper. Business AI is learning the same lesson, with much larger stakes. The Firmulate experiment shows that “reads your files before answering” isn’t a vague virtue — it’s a measurable, purchase-deciding property of an AI agent, worth exactly €55,000 in one week of one small company.
Before you let an agent near your CRM, your support queue, or your forecast, the question isn’t “does it write well?” It’s: does it finish what it starts, does it read two documents deep, and does it stay honest when someone pretends to be the boss? The full benchmark results — and the live company they came from — are public. Watch the agents work. Then decide which one you’d hire.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html