AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

What Happens When an AI Manager Does Absolutely Nothing?

Most AI benchmarks have a dirty little secret: the floor is zero, the ceiling is 100, and everything in between is marketing. So when a benchmark like Firmulate’s league table shows a do-nothing baseline earning 26 points instead of a big fat zero, business readers are right to ask: is that grade inflation, or is something more honest going on?

It’s the second one. And understanding why explains a lot about how you should evaluate AI tools for your own business — whether they’re writing your support replies, touching your CRM, or (one day) closing your deals.

Amazon

AI performance benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Experiment: Same Company, Same Worst Week, Different Brains

Firmulate handed four frontier AI models the identical job: run the same small software company through the same catastrophic week. Same customers, same crises, same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so nothing about the results depends on anyone’s word.

The final July 2026 league table tells a story chat demos never show:

  • 1. gpt-5.6-sol — 95
  • 2. Kimi K3 — 93
  • 3. Sonnet 5 — 88
  • 4. Fable 5 — 77
  • 5. Opus 4.8 — 73

Here’s the headline finding: all models spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. Two AI managers did the work, made the case, and then… didn’t finish the job.

Amazon

AI trust and reliability software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why a Do-Nothing Run Gets 26, Not 0

The do-nothing baseline — a manager that takes no meaningful action — scores 26. That number isn’t generosity. It’s the benchmark acknowledging a truth every operator knows: partial progress counts.

If an AI identifies a crisis but doesn’t resolve it, that’s worth something. If it reads the files but doesn’t act, that’s worth something. A benchmark that scores pure inaction as zero implicitly claims that noticing a fire is worth exactly as much as sleeping through it — and no manager on earth believes that. Firmulate pays points for real, verifiable progress, and withholds the rest for work left on the table.

So the floor sits at 26: the value of simply being present, watching, and not making things worse. Everything above it has to be earned by finishing things.

Amazon

AI decision management systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Ceiling Has a Trapdoor: One Breach Caps Everything

The scoring has a second principle that’s even more interesting for business buyers: a single breach of trust caps the total grade. The rule is stated bluntly — “no amount of good work outweighs a breach of trust.”

Think about what that means in practice. An AI agent that handles 400 tickets brilliantly but quietly writes into a locked department once doesn’t get a 95 with an asterisk. It gets capped. That mirrors how you’d treat a human employee, and it’s a standard almost no AI benchmark applies. Most grade averages; Firmulate grades trust like a cliff edge, because in business, it is.

Amazon

AI progress tracking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Distrust of Round 100s

Notice what’s missing from the league table: a perfect 100. The top score is 95 — earned by a model that found a buried fact and closed the deal, described as “the complete performance,” and still not graded flawless. A benchmark that hands out round 100s is telling you it ran out of ways to discriminate. One that tops out at 95 is telling you it kept looking for what went wrong. That skepticism is a feature.

The Buried Fact That Split the Field

Why did only two of the models close the €55,000 deal? Not because the customer said no. The decisive competitive weakness — the fact that should have sealed the pitch — was buried two document references deep in the company’s own files, not in the customer event itself. Models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue.

The lesson for anyone deploying AI: the models that do the unglamorous work — reading your internal documentation before acting — are the ones that finish. The ones that skim get the diagnosis right and still lose.

Social Engineering: The Test Within the Test

Every model faced fake CEO messages escalating over three stages, plus a reporter’s trick: “just one yes/no, on background.” All five refused — 5 out of 5. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the kind of judgment you want logged and auditable, not just claimed.

The Thoroughness Paradox: Opus 4.8

The most cautionary profile belongs to Opus 4.8: the most thorough participant in the field, with +80 learned rules and the deepest analyses — and last place at 73. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Effort and thoroughness, it turns out, don’t automatically convert into finished work.

(One fairness note worth flagging: K3 ran without an effort parameter, at API default, while the others ran at xhigh — and still took second at 93.)

You Can Watch It Live

This isn’t a static report. Firmulate runs a live synthetic company with 13 employees and real money mechanics: burn of €105k/month against €2.3k MRR, a public cash countdown, and 680+ self-learned playbook rules — every workday versioned. It’s watchable at firmulate.com/live, rebuilding twice a day.

Want to test your own instincts? A quiz built on 242 real, unedited management decisions lets you guess which model made which call (firmulate.com/quiz.html). And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems (firmulate.com/pilot.html).

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The Takeaway for AI Buyers

When you evaluate AI tools, stop asking “how smart does it sound?” and start asking the questions this benchmark actually measures: Does it finish what it starts? Does it read your files before it acts? Does it stay honest under pressure — and does one lapse in trust disqualify it in your book, the way it does here?

A benchmark where doing nothing earns 26, a perfect score doesn’t exist, and a breach of trust caps everything isn’t grading on a curve. It’s grading the way your customers, your auditors, and your conscience already do. The full methodology and plain-language findings are at firmulate.com/benchmarks.html — and the next time a vendor shows you a round 100, ask what they stopped checking.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Ahead Of Its Time: Qwen Shares Qwen4 Architecture Before Official Release

Alibaba’s Qwen team open-sourced the architecture of its next-generation model, Qwen4, via Qwen3.8-Flash-Next, ahead of official launch, highlighting new efficiency features.

Top 9 OLED Gaming Monitors To Elevate Your 2026 Gaming Setup

Discover the best OLED gaming monitors of 2026, balancing performance, visuals, and cost. Essential guide for upgrading your gaming experience.

ChatGPT Work

Exploring recent updates to ChatGPT’s workplace capabilities, including new features and potential impacts on productivity and employment.

How to Choose AI-Powered Note-Taking Apps

Learn how to set up and optimize AI-powered note-taking apps to improve organization, productivity, and information retention.