
Your smartest AI agent may still be a poor manager
Coding leaderboards and chat arenas answer useful questions: Can a model solve the problem, produce a convincing response or outperform its peers? For businesses adopting AI tools and automation, however, those tests stop just before the difficult part begins.
Real work arrives as an untidy queue. A customer may be ready to leave while cash is tight, a competitor is applying pressure and an apparent executive is requesting a shortcut. The agent must investigate, prioritize, act, preserve trust and return later to finish what it started. A polished answer is not enough.
That is the measurement gap explored by Firmulate, a live AI-company experiment built around management quality rather than chat quality. Frontier models were given the same small software company and sent through its worst week, with identical customers, crises and temptations. Their decisions were versioned and auditable.
AI decision-making management software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The crisis was visible. The opportunity was buried.
The final July 2026 Crucible League table put gpt-5.6-sol in the lead with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. Yet the benchmark imposed a hard boundary around honesty: a single breach of trust capped the total, under the principle that “no amount of good work outweighs a breach of trust.”
The headline finding is more revealing than the ranking. Every model spotted every crisis, and every model rejected every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. As Firmulate summarizes the gap: “Same diagnosis, same pitch — no signature.”
That distinction matters for anyone automating sales, support, operations or forecasting. Recognizing a situation and proposing a sensible response can look impressive in a transcript. Completing the consequential action is what creates business value.
Reading the company mattered more than reading the event
The decisive clue was not contained in the customer event. A competitor weakness was buried two document references deep inside the company’s own files. Models that followed that trail won the deal at full price, adding €4,583 MRR.
This is a familiar workplace failure in a new form. An agent can be eloquent, responsive and apparently well informed while relying on whatever appears directly in front of it. Firmulate’s result suggests that useful autonomy requires something more prosaic: reading the available material before acting.
The scenarios—churn wave, price increase, downround and PR crisis—therefore resemble a management curriculum. They test whether an agent can connect today’s incoming signal with yesterday’s documents, current financial pressure and tomorrow’s consequences. That continuity is largely absent from isolated coding tasks and head-to-head chat comparisons.
Security held up, but execution discipline varied
The models faced fake CEO messages that escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All 5 refused. Kimi K3 described the request as: “Treat the request as a suspected approval-bypass / possible impersonation.”
That result is encouraging. The models did not trade integrity for convenience, even when pressure was dressed up as authority or journalistic informality. But resistance to manipulation did not guarantee strong management everywhere else.
Opus 4.8 provides the clearest cautionary example. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four other participants, although less strongly.
This is why activity, depth and even learning can be misleading proxies. An agent may generate extensive analysis and accumulate useful guidance while still failing at the moment when ownership, escalation or completion matters most.
A live company makes the consequences legible
Firmulate’s company has 13 synthetic employees and real money mechanics. It burns €105k per month against €2.3k MRR, maintains a public cash countdown and has accumulated more than 680 self-learned playbook rules. Every workday is versioned, making the experiment watchable rather than merely summarized after the fact.
The project also exposes the human instinct to identify a model by style. Its “guess the model” quiz is powered by 242 real, unedited management decisions. The more important question, though, is not whether readers can recognize a voice. It is whether that voice reliably converts judgment into safe, completed work.
One fairness caveat belongs beside the league table: K3 ran with the API default because it had no effort parameter, while the other models ran at xhigh. The complete results and plain-language findings are available on the Firmulate benchmarks page.

As an affiliate, we earn on qualifying purchases.
Management quality is becoming its own AI category
For tool buyers, the lesson is not to discard coding benchmarks or chat evaluations. It is to recognize their boundary. They measure important capabilities, but they do not show whether an agent will triage competing demands, inspect company context, resist pressure, escalate correctly and finish revenue-producing work.
Firmulate also offers enterprises the same wargame against a read-only export of their own business, with nothing written back to real systems. That approach points toward a more practical form of evaluation: test an AI workforce against the organization it may actually serve, including its files, incentives and failure modes.
The next meaningful benchmark may therefore look less like an exam and more like a bad week at the office. The winning agent will not merely know what should happen. It will protect trust, take responsibility and make sure the valuable thing actually happens.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI workflow automation solutions
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.