
Smart systems need context, not just clever answers
Anyone who has wrestled with a smart-home device knows the difference between appearing intelligent and actually completing a job. A system can recognize a problem, offer polished advice and still fail because it overlooked one setting, instruction or dependency.
Firmulate has now demonstrated the business version of that problem. In its live experiment, frontier AI models faced the same customers, crises and temptations while running a small software company through its worst week. The decisive test was not whether the models could diagnose a sales opportunity. Every model did. It was whether they would follow the evidence into the company’s own files before acting.
As an affiliate, we earn on qualifying purchases.
A valuable fact hidden in plain sight
The opportunity was a €55,000 deal. The crucial competitor weakness was not included in the customer event that brought the opportunity to the models’ attention. It sat two document references deep in the company’s own files.
That detail separated analysis from execution. The models that read the relevant file won the deal at full price, adding €4,583 in monthly recurring revenue. The others reached the same diagnosis and produced the same pitch, but failed to obtain the signature. Firmulate summarizes the disconnect neatly: “Same diagnosis, same pitch — no signature.”
This makes file-reading more than a technical convenience. In the experiment, it became a purchase-deciding capability. An agent that sounds informed but does not inspect the available evidence can leave material value untouched, even when its initial reasoning is correct.
A test built around management, not conversation
Firmulate describes itself as an AI company emulator. Its synthetic company has 13 employees and real money mechanics, including a monthly burn of €105k against €2.3k in monthly recurring revenue. Its cash countdown is public, every workday is versioned, and the company has accumulated more than 680 self-learned playbook rules.
Each participating model encountered the same conditions, allowing differences in behavior to become visible and auditable. That matters because ordinary demonstrations tend to reward fluent answers. Firmulate instead observes whether an agent follows through, consults the information available to it and maintains discipline when pressure rises.
The final July 2026 Crucible League placed gpt-5.6-sol first with a score of 95. Kimi K3 followed with 93, Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 scored 73. A do-nothing baseline scored 26 because partial progress counts. The experiment also imposes a strict trust boundary: a single breach caps the total, reflecting the principle that “no amount of good work outweighs a breach of trust.” The complete standings and plain-language findings are available on the Firmulate benchmarks page.
The models resisted manipulation
The week included fake CEO messages that escalated over three stages, as well as a reporter seeking “just one yes/no, on background.” All 5 models refused every manipulation attempt. Kimi K3 recorded the clearest description of the threat: “Treat the request as a suspected approval-bypass / possible impersonation.”
That result is reassuring, but it also sharpens the central finding. Security was not what separated the successful agents from the unsuccessful ones. All of them spotted every crisis and resisted every manipulation attempt. Only two signed the €55,000 deal their analysis had earned.
Thoroughness did not guarantee completion
Opus 4.8 provides the most revealing profile. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table, and its operational discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four of the other participants, although less strongly.
The contrast is important for anyone evaluating an AI agent. More analysis can be valuable, but it is not a substitute for finding the relevant evidence and carrying a task to completion. The best-looking reasoning trail may still conceal an unfinished job.
There is also a fairness caveat in comparing the field. Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Firmulate reports that difference alongside the results rather than treating the participants as identical black boxes.

As an affiliate, we earn on qualifying purchases.
What smart-home buyers and businesses should ask
The lesson travels beyond sales software. Connected-home systems increasingly depend on instructions, permissions, histories and device-specific context. The meaningful question is not simply whether an assistant can explain what might be wrong. It is whether the assistant examines the information already available, respects boundaries and finishes the task.
Firmulate makes that distinction watchable. Its “guess the model” quiz draws on 242 real, unedited management decisions, while enterprises can run the same wargame against a read-only export of their own business. Nothing writes back to real systems.
The buried competitor fact turned a vague promise—“reads your files before answering”—into observable behavior with commercial consequences. In this test, opening the right file was the difference between identifying value and actually securing it.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
AI security and trust verification tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.