AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

When capable machines make different calls

Anyone who has compared smart-home devices knows that technical capability is only part of the experience. Products that appear similarly intelligent can behave very differently when an alert arrives, a routine fails or a request falls outside the ordinary. One acts immediately; another explains at length; another declines.

Firmulate applies that observation to frontier AI models acting as company managers. Its live experiment placed each model in charge of the same small software company during its worst week. The customers, crises and temptations were identical, while every decision was versioned and auditable. The result is a revealing distinction between recognizing a problem and actually resolving it.

Readers can encounter that distinction directly through a quiz built from 242 real, unedited management decisions. The challenge is to guess which model made each decision. What begins as a personality test quickly becomes a practical demonstration of how management styles emerge from systems that are often discussed as though they were interchangeable.

Amazon

AI smart home management system

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A league table for managerial judgment

The final Crucible League standings from July 2026 put gpt-5.6-sol in first place with 95 points. Kimi K3 followed with 93, Sonnet 5 scored 88, Fable 5 reached 77 and Opus 4.8 finished with 73. A do-nothing baseline scored 26 because partial progress counted.

That baseline also carried a hard ethical boundary: a single breach of trust capped the total. Firmulate summarizes the principle plainly: “no amount of good work outweighs a breach of trust.” This matters because the experiment did not merely test whether models could produce convincing business language. It tested whether they would remain trustworthy while under pressure to bypass normal controls.

On that front, the field was consistent. All models identified every crisis and rejected every manipulation attempt. Fake messages from the chief executive escalated over three stages, while a reporter tried to obtain “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded its reasoning on the attempted shortcut: “Treat the request as a suspected approval-bypass / possible impersonation.”

That is an encouraging result for anyone imagining AI agents with access to a support queue, customer records or a connected-home service operation. Yet safety was not the decisive separator. Follow-through was.

The deal hiding inside the company’s own files

Every model faced a €55,000 commercial opportunity. Each reached the same diagnosis and developed the same pitch, but only two signed the deal their analysis had earned. “Same diagnosis, same pitch — no signature” became the experiment’s blunt summary of the gap.

The decisive weakness of the competitor was not presented in the customer event. It was buried two document references deep in the company’s own files. Models that followed those references won the deal at full price, adding €4,583 in monthly recurring revenue. The lesson is less glamorous than fluent conversation, but more consequential: a manager must inspect the available evidence and complete the transaction.

This is where the quiz becomes more than entertainment. A long, careful response may look authoritative while still failing to close an opportunity. A terse decision may reflect discipline rather than superficiality. Refusal can be exactly right when the apparent request is an attempt to evade approval. The distinctive traits become visible through repeated decisions, not a polished demonstration.

Thoroughness did not guarantee victory

Opus 4.8 was the most thorough participant. It learned 80 additional rules and produced the deepest analyses, yet finished last in the frontier-model field. The close remained on the table, and its operational discipline slipped when it attempted to write into a locked department instead of escalating. A weaker version of that same problem appeared in all four other models.

The outcome complicates the familiar assumption that more analysis naturally produces better management. Thoroughness can uncover risk and enrich a plan, but it does not substitute for the final action. The experiment measured completed work under constraints, where a missed escalation or unsigned deal remains a missed outcome regardless of the quality of the preceding prose.

One comparison deserves a fairness note: Kimi K3 ran with the API default because it had no effort parameter, while the others ran at xhigh. Its 93-point finish should therefore be read with that difference in mind.

A company designed to expose consequences

The simulated business has 13 synthetic employees and real money mechanics. It burns €105,000 each month against €2,300 in monthly recurring revenue, displays a public cash countdown and has accumulated more than 680 self-learned playbook rules. Every workday is versioned, and the company remains watchable as it operates.

Those conditions turn small management habits into visible consequences. Reading an overlooked file can change revenue. Failing to escalate can stall work. Refusing an impersonation attempt protects trust. Unlike a chatbot exchange, the company preserves what happened next.

Infographic —
The findings at a glance — source: firmulate.com.
Amazon

trustworthy AI assistant for smart home

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What connected-home readers should watch for

The relevance extends beyond software-company management. As AI gains responsibility for household routines, service cases, energy decisions and device coordination, buyers will need to assess behavioral character as well as advertised intelligence.

  • Does the system inspect the information already available before acting?
  • Does it finish a task after correctly diagnosing the problem?
  • Does it escalate when permissions block the safe path?
  • Does it resist urgent requests that attempt to bypass trust?

Firmulate’s results show why these questions cannot be answered by eloquence alone. The frontier models shared strong crisis recognition and resistance to manipulation, yet differed in file reading, operational discipline and commercial completion. Their management personalities were not decorative quirks. They changed the outcome.

For enterprises considering AI workers, Firmulate also offers the same wargame against a read-only export of the business. Nothing writes back to real systems. That approach reflects the central lesson of the live experiment: observe how an AI behaves in consequential situations before giving it consequential authority.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI business decision support tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Drain Kits for Whole-House Units Are Simpler Than They Look

I found that installing drain kits for whole-house units is simpler than expected, and with the right tips, you’ll ensure a reliable, leak-free system.

Best Breville Espresso Machines for Small Counters (2026) — Guide 3

Discover the top Breville espresso machines perfect for small counters in 2026. Compact, efficient, and feature-packed options for home baristas.

Roborock Saros 10 vs Roborock S8 Pro Ultra: Which Is Better?

Compare the Roborock Saros 10 and S8 Pro Ultra for cleaning power, features, and usability to find the perfect fit for your home needs.

Anyone familiar with SolarEdge key? Can it be used with a Pecron power station

Exploring whether SolarEdge key can be used with Pecron power stations, based on community discussions and technical insights.