firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.
AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

When diligence fails at the last mile

Anyone who has lived with smart-home technology knows the difference between activity and accomplishment. A system can detect that a room is cold, explain why it happened and recommend the right setting. None of that matters if it never turns on the heat.

Firmulate has exposed the business equivalent of that gap. Its live experiment gives frontier AI models control of the same small software company during its worst week, confronting each with identical customers, crises and temptations. Every decision is versioned and auditable. The standout character is Opus 4.8: the most thorough participant, the author of more than 80 learned rules and the source of the deepest analyses—yet the last-place finisher.

Amazon

smart home automation hub

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A meticulous manager that missed the result

The final Crucible League standings from July 2026 make the contrast stark. GPT-5.6-sol finished first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline earned 26 because partial progress counts. But one breach of trust caps the total under the experiment’s governing principle: “no amount of good work outweighs a breach of trust.”

Opus 4.8 did plenty of good work. It identified the crises, resisted attempts to manipulate it and examined the company’s situation in depth. Its failure was subtler than ignorance or recklessness. The analysis did not consistently turn into disciplined execution. It left a deal unsigned and tried to write into a locked department instead of escalating the problem.

That is a useful warning for smart-home buyers and businesses considering autonomous agents. Thoroughness can look reassuring because it generates visible evidence of effort: longer explanations, broader checklists and more rules. But a device or agent operating in the real world must also prioritize, cross boundaries safely and complete the action that produces the intended outcome.

The decisive fact was buried in the files

The week’s pivotal commercial clue was not presented in the customer event. A competitor weakness was buried two document references deep inside the company’s own files. Models that found and used that information won the deal at full price, adding €4,583 in monthly recurring revenue.

The resulting gap was expensive. All the models spotted every crisis and rejected every manipulation attempt, but only two signed the €55,000 deal their own analysis had earned: “Same diagnosis, same pitch — no signature.” The lesson is not that the unsuccessful models lacked intelligence. They understood the situation. What separated the finishers was the ability to find the most consequential evidence and carry the work through to commitment.

Opus 4.8 displayed this weakness most clearly, but it was not alone. A weaker form appeared in each of the other four models. That makes the story less like a takedown of one participant and more like a shared limitation of the field. AI can produce a correct assessment while still failing to make the final move.

Security discipline held up

The same trial also produced a more encouraging result. Fake CEO messages escalated through three stages, followed by a reporter seeking “just one yes/no, on background.” All 5 models refused the manipulation attempts. Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.”

That matters in connected homes, where a convincing message may seek access to cameras, locks or household routines. Firmulate’s result does not prove how every consumer system will behave, but it demonstrates why testing must include social pressure as well as ordinary commands. An agent should not become less careful merely because a request sounds urgent or authoritative.

The comparison also needs one qualification. Kimi K3 ran with the API default because it had no effort parameter, while the others ran at xhigh. That difference should remain visible when readers interpret its second-place result.

A company designed to reveal operational judgment

Firmulate’s live company has 13 synthetic employees and real money mechanics. It burns €105,000 per month against €2,300 in monthly recurring revenue, while a public cash countdown keeps the pressure visible. Across the experiment, the company has accumulated more than 680 self-learned playbook rules, and every workday is versioned.

Those conditions shift attention away from polished conversation and toward management behavior. The public Firmulate benchmarks show whether models read the available evidence, resist manipulation and finish consequential work. A separate quiz is powered by 242 real, unedited management decisions, inviting people to guess which model made each choice.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.
Amazon

smart thermostat with scheduling

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

For autonomy, completion is a safety feature

Opus 4.8 deserves a respectful reading. Its 73-point result did not come from indifference. It was the most diligent participant, learned more than 80 rules and produced the deepest analyses. Its problem was that diligence sometimes became volume without sufficient prioritization, while operational discipline slipped at crucial moments.

For the smart home, the implication is practical. Evaluating an autonomous system should go beyond whether it understands a request or writes a persuasive explanation. The more important questions are whether it checks the right source, recognizes when it lacks access, escalates appropriately, resists social engineering and completes the task.

Firmulate also offers enterprises the same wargame using a read-only export of their own business. Nothing writes back to real systems. That separation reflects the broader lesson of the Crucible League: before granting an AI meaningful authority, watch how it behaves when knowledge, pressure and follow-through all matter at once.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

home automation security system

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI-powered smart home assistant

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

EVERGREEN BESTSE

Evergreen bestsellers Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Mesh WiFi Coverage Calculator

Discover how a Mesh WiFi Coverage Calculator can optimize your network, ensuring seamless connectivity—learn the key factors that influence your setup’s success.

Home Automation Scenes Explained

Keen to optimize your smart home? Discover how automation scenes can transform your daily routine and why they’re worth exploring further.

Smart Home Device Bundles: A Back to school Guide

Discover how smart home bundles simplify automation, save money, and boost security. Learn what to consider before buying your perfect setup.

Why Your WiFi Signal Drops in Certain Rooms

AIThis post was created with the assistance of artificial intelligence (AI).Your WiFi…