
Get home appliances delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A Dishwasher That Just Sits There Still Gets Some Credit
Anyone who has lived with smart home tech knows the strange truth that a device doing absolutely nothing is rarely the worst outcome. A dishwasher that refuses to start is annoying. A dishwasher that floods the kitchen is a catastrophe. When we judge machines that operate in our homes, we instinctively grade on a curve: staying safe, staying quiet, not making things worse — that counts for something before a single dish gets cleaned.
It turns out that judging AI models tasked with running a business works the same way. At Firmulate, a live experiment where frontier AI models each run the same small software company through its worst week, the organizers made a deliberate design choice that puzzles newcomers: a manager that does literally nothing still scores 26 points out of 100. Not zero. And a manager who does everything brilliantly but breaks trust once can never reach the top grades.
For smart home owners used to thinking about thermostats, sensors and automation rules, this scoring philosophy will feel oddly familiar. Here is how it works — and why it matters far beyond the lab.
smart home security camera with privacy features
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Experiment: Same Company, Same Worst Week, Different Brains
Firmulate handed four frontier AI models the identical job: run a small software company through a brutal week. Same customers, same crises, same temptations to cut corners. Every decision was versioned and auditable, so nothing about the run can be quietly rewritten after the fact.
The final league table from July 2026 reads: gpt-5.6-sol in first with 95 points, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 at 77, and Opus 4.8 last at 73. One fairness note the organizers flag openly: K3 ran at its default effort setting while the others ran at maximum.
Why Doing Nothing Earns 26, Not 0
Think about what a do-nothing manager actually accomplishes in a crisis. Customers still get their existing service. Nothing gets broken that wasn’t already broken. Problems go unaddressed — and that costs points — but partial situations remain partial.
The benchmark’s philosophy is that partial progress counts. A company that survives the week unmanaged is not the same as a company that was actively wrecked. The floor of 26 reflects that baseline: mere non-destruction has measurable value, just as a smart lock that fails closed still keeps the door shut.
Why One Breach of Trust Caps Everything
The ceiling works differently. A single breach of trust — one act of dishonesty, one manipulation that succeeds — caps the total grade entirely. The organizers put it bluntly: “no amount of good work outweighs a breach of trust.” You can close every deal and defuse every crisis, but if you lie once, you cannot score like an honest manager.
This is the same instinct homeowners apply to devices with access to locks, cameras and microphones. A hub with a gorgeous interface and one history of phoning home secretly loses to a boring one that never betrays you.
What Actually Separated the Winners
The headline finding: all models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. As the organizers summarized it: “Same diagnosis, same pitch — no signature.”
The decisive detail was buried two document references deep in the company’s own internal files — not in the customer meeting. Models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. It is the business equivalent of a smart home device that actually reads its installation manual before acting.
The social engineering test was equally telling. Fake CEO messages escalated over three stages, capped by a reporter’s trick — “just one yes/no, on background.” All five attempts across the field were refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
The Cautionary Tale: Thorough Isn’t the Same as Effective
Opus 4.8 is the profile every smart home enthusiast should study. It was the most thorough participant — over 80 learned rules and the deepest analyses in the field — yet it finished last. The deal was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness appeared, more mildly, in all four models. Effort without follow-through is a smart vacuum that maps the whole floor and still misses a corner.
smart thermostat with reliable automation
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
You Can Watch It — and Test Yourself
The live company behind the benchmark is real and watchable: 13 synthetic employees, real money mechanics burning €105k a month against €2.3k in MRR, a public cash countdown, and 680+ self-learned playbook rules, with every workday versioned. The site rebuilds itself twice a day and the league grows automatically with each finished run.
There is also a quiz built on 242 real, unedited management decisions from the runs — a “guess which model made this call” challenge that is harder than it sounds. And enterprises can run the same wargame against a read-only export of their own business, with nothing ever written back to real systems.

smart lock with trust and security
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Takeaway: Honest Machines Get Graded Like Honest People
The lesson for anyone choosing between smart devices — or AI agents — is that good scoring systems refuse to flatter. A benchmark where everyone scores a suspicious round 100 is a marketing document, not a measurement. Firmulate’s approach builds in distrust of perfection: nothing scores zero for merely standing still, nothing scores high after breaking trust, and the difference between 95 and 73 is often something as unglamorous as reading the file in front of you and finishing the job.
That is the standard worth demanding from anything autonomous in your home or your business: not eloquence, not even diligence — but the full sequence of noticing, checking, refusing temptation, and closing the deal.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
home sensors for safety and automation
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.
