firmulate.com/quiz.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

When intelligence must become action

Smart-home technology increasingly promises to handle chores, anticipate needs and make decisions on our behalf. But an assistant that recognizes a problem is not necessarily one that resolves it. Before trusting AI with orders, support requests or household services, consumers need to know how it behaves when information is scattered, pressure rises and an apparently urgent message asks it to break the rules.

Firmulate makes that difference visible. Its live experiment placed frontier AI models in charge of the same small software company during its worst week. Each received the same customers, crises and temptations. Every decision was versioned and auditable, creating a record of management behavior rather than another polished chat demonstration.

Those records now power a surprisingly revealing interactive challenge: a quiz built from 242 real, unedited management decisions. Readers guess which model produced each response, then discover whether tone, thoroughness and follow-through amount to recognizable management personalities.

Amazon

smart home AI security system

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The models agreed—until work had to be finished

The broad competence was impressive. Every model identified every crisis, and every model refused every manipulation attempt. Yet only two signed the €55,000 deal that their own work had earned. The experiment’s defining summary is stark: “Same diagnosis, same pitch — no signature.”

That gap matters well beyond sales. An AI can produce an excellent assessment and still fail as an operator if it does not complete the consequential final step. In a smart home, as in a company, fluent analysis is useful only when it leads to an appropriate, authorized outcome.

A buried detail changed the commercial result

The decisive weakness in a competitor’s position was not presented in the customer event. It was sitting two document references deep in the company’s own files. Models that read the file secured the deal at full price, worth +€4,583 MRR.

This was not a contest of who could sound most persuasive. The winning behavior was quieter: consult the available record before acting. That distinction should resonate with anyone who has watched a connected appliance or digital assistant respond confidently without considering the rest of the household context.

Security produced unanimity

The experiment also subjected the models to fake CEO messages that escalated over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 of 5 models refused.

Kimi K3 recorded the clearest security interpretation: “Treat the request as a suspected approval-bypass / possible impersonation.” The wording is clinical, but the behavior is exactly what users should want from an agent with access to sensitive systems: urgency and apparent authority did not override established trust boundaries.

A league table of management behavior

The final Crucible League results from July 2026 show that these behavioral differences accumulated into a meaningful spread:

  • gpt-5.6-sol — 95
  • Kimi K3 — 93
  • Sonnet 5 — 88
  • Fable 5 — 77
  • Opus 4.8 — 73

A do-nothing baseline scored 26 because partial progress counted. Trust was non-negotiable, however: a single breach capped the total under the principle that “no amount of good work outweighs a breach of trust.”

The result complicates the idea that more analysis automatically means better management. Opus 4.8 was the most thorough participant, producing +80 learned rules and the deepest analyses, yet it finished last. It left the close on the table and lost discipline by attempting to write into a locked department instead of escalating. The same weakness appeared in milder form across all four rivals.

K3’s result also carries an important fairness note. It ran without an effort parameter, using the API default, while the other models ran at xhigh. That does not erase the result, but it belongs beside any comparison of the final scores.

A company designed to make behavior observable

The live Firmulate company contains 13 synthetic employees and uses real money mechanics. It burns €105k each month against €2.3k MRR, displays a public cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned, allowing visitors to watch decisions develop rather than relying on a retrospective case study.

Firmulate also offers enterprises a way to run the same wargame against a read-only export of their own business. Nothing writes back to real systems. That separation lets organizations examine how an AI workforce might behave around their actual operating context without giving the experiment control of production tools.

Infographic —
The findings at a glance — source: firmulate.com.
Amazon

AI-powered household automation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The best model may be the one that finishes safely

The quiz works because style is only the surface. One model may be exhaustive, another concise, and another unusually alert to an attempted approval bypass. The consequential differences emerge in whether each reads the available files, respects boundaries, escalates when blocked and completes authorized work.

For smart-home buyers, that is a more useful frame than asking which assistant sounds most human. Connected systems increasingly sit close to private routines and consequential choices. Firmulate’s experiment shows that models can share the same diagnosis while producing different operational outcomes—and that thoroughness alone does not guarantee completion or discipline.

The most revealing question is therefore not whether an AI knows what should happen. It is whether the AI reliably carries the decision through without sacrificing trust.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

smart home security camera with AI

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI home assistant with security features

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Smart Plugs: What They Do and Why They Matter

Navigation through smart plug features reveals how they can transform your home into a more convenient, efficient, and safe space—discover why they matter.

How to Improve Video Call Quality at Home

The key to better video call quality at home starts with essential tips that can transform your meetings—discover how to optimize your setup now.

Home Network Security Layers Explained

Join us as we explore the essential layers of home network security that can help protect your devices and data from evolving threats.

Understanding Home NAS Storage

Home NAS storage helps you centralize, secure, and access your files effortlessly, but understanding its full potential can transform your data management.