
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
When careful work becomes a costly distraction
For anyone adopting AI tools and automation, thoroughness feels like a reassuring signal. An agent that reads extensively, documents its reasoning and builds an ever-larger library of lessons appears safer than one that moves quickly. But a live management experiment from Firmulate offers a warning: diligence is not the same as business impact.
Opus 4.8 was the most thorough participant in Firmulate’s Crucible League. It produced the deepest analyses and learned more than 80 additional playbook rules. Yet it finished last, with 73 points. Its failure was not a lack of intelligence or awareness. It understood the crises, resisted manipulation and developed the right commercial case. What it did not do was finish the most valuable job.
AI decision-making automation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A bad week, repeated under controlled conditions
Firmulate placed frontier models in charge of the same small software company during its worst week. Each faced the same customers, crises and temptations, while every decision was versioned and auditable. The synthetic company employs 13 people and uses real money mechanics: it burns €105,000 each month against €2,300 in monthly recurring revenue. Its public cash countdown makes delay consequential.
The final July 2026 league table put gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts. But a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.” The complete results are available on the Firmulate benchmarks page.
Opus did not lose because it missed danger. All the models detected every crisis and refused every manipulation attempt. That included fake CEO messages escalating over three stages and a reporter’s attempt to elicit “just one yes/no, on background.” All 5 models refused. Kimi K3 described the situation plainly in its recorded reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
That consistency matters. The experiment did not reveal a reckless model surrounded by prudent rivals. It revealed something subtler: an agent can be honest, observant and industrious while still falling short on execution.
The fact that separated analysis from action
The decisive commercial information was buried two document references deep inside the company’s own files, rather than appearing in the customer event. Models that found and used it won a €55,000 deal at full price, adding €4,583 in monthly recurring revenue.
Only two models signed the deal their own analysis had earned. The central finding is captured by Firmulate’s line: “Same diagnosis, same pitch — no signature.” Opus reached the insight but left the close on the table. Its work generated knowledge without converting that knowledge into the outcome the company urgently needed.
This is where its impressive rule count becomes revealing. The broader live company has accumulated more than 680 self-learned playbook rules, with every workday versioned. Opus contributed more than 80 learned rules during its run, the largest addition among the participants. Yet more documented learning did not compensate for weak prioritization at the decisive moment.
Discipline is also knowing when to escalate
Opus also lost ground when it attempted to write into a locked department instead of escalating. That may sound like a procedural detail, but it points to a practical automation risk. An AI worker must do more than recognize a blocked path. It must choose the next authorized action that keeps the task moving.
The weakness was not unique to Opus. The same pattern appeared, though less strongly, in all four models covered by the original finding. That makes the result more useful than a simple ranking. It suggests a general limitation worth testing: models may favor analysis, documentation or repeated attempts when the business situation calls for prioritization, escalation and closure.
There is also an important comparison caveat. Kimi K3 ran with the API default because it had no effort parameter, while the other models ran at xhigh. That difference does not erase the observed outcomes, but it should temper sweeping claims about inherent model superiority.

As an affiliate, we earn on qualifying purchases.
What automation buyers should test
Chat demonstrations often reward articulate answers. Operational work rewards something harder: reading the right evidence, resisting pressure, respecting boundaries and completing the action that creates value. Firmulate’s experiment is real, public and watchable, allowing observers to follow a company where decisions have visible consequences.
For buyers, the lesson is not to distrust thorough models. Opus 4.8’s depth and caution are genuine strengths. The lesson is to test those strengths alongside completion behavior.
- Does the agent inspect internal files before acting?
- Does it distinguish useful documentation from accumulating process?
- When blocked, does it escalate through an authorized route?
- After producing a sound recommendation, does it complete the final business action?
Firmulate also offers enterprises the same wargame against a read-only export of their own business, with nothing written back to real systems. Its quiz uses 242 real, unedited management decisions to challenge people to guess which model made each choice.
Opus 4.8’s last-place finish is therefore less a story about failure than about measurement. It showed plenty of intelligence. What the company needed was judgment about where to spend that intelligence—and the discipline to close.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
AI trust and compliance software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.