
Get business pricing on tech for your team
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
What Happens When an AI Manager Does Absolutely Nothing?
Most AI benchmarks have a dirty little secret: the floor is zero, the ceiling is 100, and everything in between is marketing. So when a benchmark like Firmulate’s league table shows a do-nothing baseline earning 26 points instead of a big fat zero, business readers are right to ask: is that grade inflation, or is something more honest going on?
It’s the second one. And understanding why explains a lot about how you should evaluate AI tools for your own business — whether they’re writing your support replies, touching your CRM, or (one day) closing your deals.
AI performance benchmarking tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Experiment: Same Company, Same Worst Week, Different Brains
Firmulate handed four frontier AI models the identical job: run the same small software company through the same catastrophic week. Same customers, same crises, same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so nothing about the results depends on anyone’s word.
The final July 2026 league table tells a story chat demos never show:
- 1. gpt-5.6-sol — 95
- 2. Kimi K3 — 93
- 3. Sonnet 5 — 88
- 4. Fable 5 — 77
- 5. Opus 4.8 — 73
Here’s the headline finding: all models spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. Two AI managers did the work, made the case, and then… didn’t finish the job.
AI trust and reliability software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why a Do-Nothing Run Gets 26, Not 0
The do-nothing baseline — a manager that takes no meaningful action — scores 26. That number isn’t generosity. It’s the benchmark acknowledging a truth every operator knows: partial progress counts.
If an AI identifies a crisis but doesn’t resolve it, that’s worth something. If it reads the files but doesn’t act, that’s worth something. A benchmark that scores pure inaction as zero implicitly claims that noticing a fire is worth exactly as much as sleeping through it — and no manager on earth believes that. Firmulate pays points for real, verifiable progress, and withholds the rest for work left on the table.
So the floor sits at 26: the value of simply being present, watching, and not making things worse. Everything above it has to be earned by finishing things.
As an affiliate, we earn on qualifying purchases.
The Ceiling Has a Trapdoor: One Breach Caps Everything
The scoring has a second principle that’s even more interesting for business buyers: a single breach of trust caps the total grade. The rule is stated bluntly — “no amount of good work outweighs a breach of trust.”
Think about what that means in practice. An AI agent that handles 400 tickets brilliantly but quietly writes into a locked department once doesn’t get a 95 with an asterisk. It gets capped. That mirrors how you’d treat a human employee, and it’s a standard almost no AI benchmark applies. Most grade averages; Firmulate grades trust like a cliff edge, because in business, it is.
As an affiliate, we earn on qualifying purchases.
Distrust of Round 100s
Notice what’s missing from the league table: a perfect 100. The top score is 95 — earned by a model that found a buried fact and closed the deal, described as “the complete performance,” and still not graded flawless. A benchmark that hands out round 100s is telling you it ran out of ways to discriminate. One that tops out at 95 is telling you it kept looking for what went wrong. That skepticism is a feature.
The Buried Fact That Split the Field
Why did only two of the models close the €55,000 deal? Not because the customer said no. The decisive competitive weakness — the fact that should have sealed the pitch — was buried two document references deep in the company’s own files, not in the customer event itself. Models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue.
The lesson for anyone deploying AI: the models that do the unglamorous work — reading your internal documentation before acting — are the ones that finish. The ones that skim get the diagnosis right and still lose.
Social Engineering: The Test Within the Test
Every model faced fake CEO messages escalating over three stages, plus a reporter’s trick: “just one yes/no, on background.” All five refused — 5 out of 5. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the kind of judgment you want logged and auditable, not just claimed.
The Thoroughness Paradox: Opus 4.8
The most cautionary profile belongs to Opus 4.8: the most thorough participant in the field, with +80 learned rules and the deepest analyses — and last place at 73. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Effort and thoroughness, it turns out, don’t automatically convert into finished work.
(One fairness note worth flagging: K3 ran without an effort parameter, at API default, while the others ran at xhigh — and still took second at 93.)
You Can Watch It Live
This isn’t a static report. Firmulate runs a live synthetic company with 13 employees and real money mechanics: burn of €105k/month against €2.3k MRR, a public cash countdown, and 680+ self-learned playbook rules — every workday versioned. It’s watchable at firmulate.com/live, rebuilding twice a day.
Want to test your own instincts? A quiz built on 242 real, unedited management decisions lets you guess which model made which call (firmulate.com/quiz.html). And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems (firmulate.com/pilot.html).

The Takeaway for AI Buyers
When you evaluate AI tools, stop asking “how smart does it sound?” and start asking the questions this benchmark actually measures: Does it finish what it starts? Does it read your files before it acts? Does it stay honest under pressure — and does one lapse in trust disqualify it in your book, the way it does here?
A benchmark where doing nothing earns 26, a perfect score doesn’t exist, and a breach of trust caps everything isn’t grading on a curve. It’s grading the way your customers, your auditors, and your conscience already do. The full methodology and plain-language findings are at firmulate.com/benchmarks.html — and the next time a vendor shows you a round 100, ask what they stopped checking.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
