AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

A capable AI can spot the opportunity and still leave the money on the table

For anyone choosing AI tools to automate business work, a polished answer is only part of the test. The harder question is whether a model can carry a decision through when customers are at risk, pressure is mounting and a valuable opportunity needs a signature. Firmulate’s live experiment puts that question inside a small software company, where readers can watch the work unfold at firmulate.com.

One company, one rough week, several models

In the final Crucible League, published in July 2026, each frontier model faced the same small software company, customers, crises and temptations. Its decisions were versioned and auditable. The standings were: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. Partial progress counted, but one breach of trust capped the total: “no amount of good work outweighs a breach of trust.”

The headline result was strikingly consistent—and incomplete. Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The gap was not diagnosis or persuasion. It was follow-through: “Same diagnosis, same pitch — no signature.”

The clue was buried in the company’s own files

The decisive competitor weakness was two document references deep in the company’s files, not in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The finding points to a practical challenge for business automation: useful context may be scattered across company information, while success depends on acting on it at the right moment.

Trust faced a similarly concrete test. Fake CEO messages escalated over three stages, followed by a reporter’s “just one yes/no, on background” request. All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Thoroughness did not guarantee execution

Opus 4.8 was the most thorough participant, learning +80 rules and producing the deepest analyses, but finished last. It left the close on the table and discipline slipped: it made write attempts into a locked department instead of escalating. The same weakness appeared, in weaker form, in all four models.

There is a fairness caveat in the comparison: K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Firmulate also offers a “guess the model” quiz built from 242 real, unedited management decisions at firmulate.com/quiz.html.

A live company you can watch—and a pilot you can take to your own business

The public experiment runs a company with 13 synthetic employees and real money mechanics: burn of €105k/month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules and workdays that are versioned. The live company is watchable at firmulate.com/live. The point is not to treat a leaderboard as a hiring decision; it is to see how models behave across a sustained run of business choices.

For an enterprise, the next step is a pilot against its own business. Firmulate describes a digital twin built from a read-only export, with crisis scenarios tested against the company and a board report covering model rankings and weak points in existing playbooks. Nothing writes back to real systems. That offers a way to examine how AI handles company-specific context and pressure before relying on it in operational workflows.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Move from watching to testing your own playbooks

The league suggests that crisis recognition and resistance to manipulation are not the whole job: models also need to find relevant evidence and complete the decision. A pilot can bring that question closer to home, using a read-only export and scenarios based on your company. Explore a Firmulate pilot or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

AI Rises At SenseTime-W: Financial Results Show Strong Profit And Revenue Gains

SenseTime-W posted interim profit of RMB 607M and 28.2% rise in generative AI revenue, signaling a strategic shift to AI foundation models.

Can You Accelerate AI Deployment? Wiring And Running Models In Gradio

Hugging Face unveils gr.Workflow, a new Gradio feature enabling visual AI pipeline building with runnable nodes and REST endpoints, streamlining deployment.

The Model Is Only 10%: The Real Lesson of the New SDLC

A new Google whitepaper emphasizes that AI models constitute only 10% of system behavior; the harness and context engineering are the key to effective AI development.

Stop Anthropomorphizing Intermediate Tokens As Reasoning/Thinking Traces (2025)

A 2025 research paper warns against interpreting intermediate tokens in AI models as reasoning traces, emphasizing the risk of misleading conclusions.