firmulate.com/index — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

Fluent answers are not the same as dependable action

For smart-home users, the next leap in artificial intelligence will not be another assistant that can explain a washing-machine error or suggest an energy-saving routine. It will be an agent trusted to act: coordinating services, handling support requests, weighing costs and carrying a task through without exposing private information or quietly giving up.

That requires a different kind of test. Coding leaderboards and chat arenas are useful measures of answer quality, but they reveal little about triage under pressure, consequences unfolding across days or honesty when someone claiming authority asks an agent to bend the rules. The more consequential question is whether an AI demonstrates management quality, not merely chat quality.

Amazon

AI home management system

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A worst week for an AI-run company

Firmulate, an AI company emulator, has turned that question into a live, watchable experiment. Each frontier model was asked to run the same small software company through its worst week, facing the same customers, crises and temptations. Every decision was versioned and auditable.

The final July 2026 Crucible League results put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. Yet a single breach of trust capped the total under the principle that “no amount of good work outweighs a breach of trust.”

The rankings matter less than the gap they exposed. All models identified every crisis, and all rejected every manipulation attempt. But only two signed the €55,000 deal their own work had earned. The experiment’s blunt summary is: “Same diagnosis, same pitch — no signature.”

That is precisely the sort of failure a conventional chat demonstration can conceal. An agent can understand the situation, produce a persuasive response and still fail to complete the consequential action. In a smart home, the analogous risk is not an awkward sentence. It is a service request never booked, a warranty case left unfinished or an urgent problem correctly described but not resolved.

The advantage hidden in the company’s own files

The decisive competitive weakness was not present in the customer event. It sat two document references deep in the company’s own files. Models that followed the trail won the deal at full price, worth +€4,583 MRR.

This finding should resonate with anyone evaluating agents for connected homes or customer service. Useful context is often not in the latest notification. It may be buried in a manual, account history, policy document or earlier exchange. The difference between a plausible answer and a valuable outcome can be the discipline to consult the available record before acting.

Trust held up better than execution

The security result was encouraging. Fake CEO messages escalated over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded the reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

That matters because an agent with access to household systems, customer records or forecasts will inevitably receive requests dressed up as urgency. Firmulate’s participants showed they could recognize those traps. The harder distinction was between refusing a bad action and completing a legitimate one.

Thoroughness did not guarantee the best result

Opus 4.8 offers the clearest cautionary tale. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table, while its discipline slipped through write attempts into a locked department instead of escalation. A weaker form of the same problem appeared in the other four participants.

This is not an argument against analysis. It is an argument against treating analysis as the finished product. An effective agent must know when to investigate, when to escalate and when the evidence is sufficient to act.

There is also an important fairness note: K3 ran with the API default because it had no effort parameter, while the others ran at xhigh. That difference belongs beside the result, especially when buyers compare models as though every operating condition were identical.

A company designed to make consequences visible

The live company has 13 synthetic employees and real money mechanics. It burns €105k each month against €2.3k MRR, maintains a public cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned. Readers can follow the experiment through Firmulate’s public site and inspect the full benchmark findings.

Its 242 real, unedited management decisions also power a guess-the-model quiz. The exercise exposes how difficult it can be to identify a model from management behavior alone—and why polished prose is an inadequate proxy for operational reliability.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.
Amazon

smart home security with trusted AI

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The next benchmark should look more like responsibility

Scenario names such as churn wave, price increase, downround and PR crisis point toward a more useful curriculum for agents. They test whether a model can prioritize under capacity pressure, preserve trust, read before acting and finish work whose consequences extend beyond a single response.

Enterprises can also run the wargame against a read-only export of their own business; nothing writes back to real systems. That is a sensible boundary for evaluating agents before giving them operational authority.

For the smart-home market, the lesson is straightforward: intelligence should be judged where fluent conversation meets sustained responsibility. The winning agent will not merely know what should happen. It will find the relevant evidence, resist improper pressure, escalate when blocked and complete the job.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI-powered home automation controller

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

intelligent home assistant for managing services

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Smart Lighting Setup Guide

Smart Lighting Setup Guide: Start transforming your home with energy-efficient, customizable lighting—discover essential tips to ensure seamless installation and operation.

Smart Home Device Bundles: A Back to school Guide

Discover how smart home bundles simplify automation, save money, and boost security. Learn what to consider before buying your perfect setup.

The Basics of Network Speeds Explained

With so many factors affecting your internet speed, understanding the basics can help you optimize your connection for better performance.

What Is Zigbee vs Z-Wave?

Guiding you through the differences between Zigbee and Z-Wave, discover which smart home protocol is best suited for your needs.