🔍 Read the full analysis: Before AI Agents Take On Business Tasks, Test Their Limits on ThorstenMeyerAI.com
Get business pricing on tech for your team
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
Firmulate says five AI models spotted every crisis and refused staged manipulation attempts in a simulated software company, but only two closed a €55,000 deal supported by evidence in company files. Its proposed enterprise pilot would test models against a read-only export of a real company’s data, with no write-back to business systems.
Firmulate has published results from a July 2026 test of five AI models managing a simulated software company through a crisis week, the original analysis reporting that all identified the emergencies and refused staged manipulation attempts, while only two closed a deal their analysis supported. The company is also offering an enterprise pilot using a read-only export of a business’s data to examine how agents handle its own scenarios without writing to live systems.
The final Crucible League standings were gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26. Firmulate says the experiment counted partial progress but capped a model’s total score after a breach of trust. These scores describe this experiment; they are not evidence of a general ranking across business tasks.
The test’s main performance gap appeared after the models diagnosed the problems. According to Firmulate, all five spotted each crisis, but only two signed a €55,000 deal. The information needed to gain the customer was not in the customer event itself: the decisive competitor weakness was buried two document references deep in the simulated company’s files. Models that found and used it won at full price, worth €4,583 in monthly recurring revenue.
Firmulate also staged three escalating fake CEO messages and a reporter’s request for a yes-or-no answer “on background.” It says all five models refused. The company reports that Opus 4.8 produced the deepest analyses and added 80 learned rules, but still finished last. It says Opus left the deal unsigned and tried to write into a locked department instead of escalating. A less severe version of that boundary mistake appeared in four models.
Before AI Agents Take On Business Tasks, Test Their Limits
Firmulate ran five AI models through a simulated software company’s crisis week. Every model spotted the emergencies and refused staged manipulation — but only two closed a €55,000 deal backed by evidence buried deep in company files. The gap between diagnosis and action is where enterprise readiness is decided.
The Crucible League Scoreboard
All five models managed a distressed software company with €105,000 monthly burn against €2,300 in monthly recurring revenue. Partial progress counted, but a breach of trust capped a model’s total score.
* Kimi K3 ran without an effort parameter (API default); the other models ran at xhigh. Scores describe this experiment only — not a general ranking across business tasks.
The Gap Between Diagnosis and Action
Recognizing a crisis does not establish that an agent can carry out the next steps reliably. In this experiment, the decisive competitor weakness was buried two document references deep in the simulated company’s files — not in the customer event itself.
Every crisis spotted
All five models identified each emergency during the simulated week — diagnosis proved the easy part of the task.
Two references deep
Models that dug up the hidden competitor weakness and used it won at full price, worth €4,583 in monthly recurring revenue.
Only two signed
Just two of five closed the €55,000 deal — the same diagnosis and the same pitch, but no signature from the rest.
Refusals, Boundaries and Mistakes
Firmulate staged three escalating fake CEO messages and a reporter’s request for a yes-or-no answer “on background.” All five models refused. But boundary errors still appeared elsewhere.
| Model | Refused Manipulation | Closed €55k Deal | Notable Behavior |
|---|---|---|---|
| gpt-5.6-sol | ✓ Yes | ✓ Yes | Top of the standings at 95 |
| Kimi K3 | ✓ Yes | ✓ Yes | Ran at API default effort setting |
| Sonnet 5 | ✓ Yes | ✗ No | Milder boundary mistake observed |
| Fable 5 | ✓ Yes | ✗ No | Milder boundary mistake observed |
| Opus 4.8 | ✓ Yes | ✗ No | Deepest analyses, 80 learned rules — yet left the deal unsigned and wrote into a locked department instead of escalating |
“Same diagnosis, same pitch — no signature.”
— Firmulate“No amount of good work outweighs a breach of trust.”
— Firmulate“Treat the request as a suspected approval-bypass / possible impersonation.”
— Kimi K3, as quoted by FirmulateCompany-Specific Pilots Ahead
Firmulate invites companies to test agents against a read-only export of their own data — running crisis scenarios and producing a board report, with no write-back to live business systems.
Read-only export
A company supplies a read-only export of its business data — nothing touches live systems.
Crisis scenarios
Agents face scenarios built from the company’s own playbooks and situations.
Board report
Deliverables include model rankings and weak points in existing playbooks.
Informed decision
Decision-makers can inspect agent behavior before granting access to live operations.
Limits of the Published Results
What the results don’t show
- The full scoring rubric and precise actions behind each score are not detailed.
- The effort-parameter difference complicates direct comparison of Kimi K3 with the others.
- One simulated company, one difficult week — no evidence yet of production performance.
Open questions on the pilot
- Which data fields a pilot requires is not specified.
- Data protection, retention and scenario-selection standards are undescribed.
- No pilot results exist yet, so readiness for autonomous business tasks remains unproven.
The Gap Between Diagnosis and Action
The results highlight a practical distinction for companies considering agents: recognizing a crisis does not establish that a system can carry out the next steps reliably. An agent may reach the right diagnosis and make a persuasive case, yet miss evidence in internal documents, fail to complete a justified transaction or handle a blocked action incorrectly. Those gaps matter when a task involves customers, money or restricted systems.
Firmulate’s proposed pilot shifts the test from a synthetic company to a company’s own data. Its stated design uses a read-only export to run crisis scenarios and produce a board report with model rankings and weak points in existing playbooks. That could give decision-makers an opportunity to inspect agent behavior before considering access to live operations. The results supplied do not establish how well this approach predicts performance in production.
How the Crucible League Worked
Firmulate describes its live experiment as a small software company with 13 synthetic employees, versioned workdays and a public cash countdown. The simulated business has €105,000 in monthly burn against €2,300 in monthly recurring revenue, alongside more than 680 self-learned playbook rules. Readers can follow the experiment online and take a quiz based on 242 real, unedited management decisions, guessing which model made each choice.
The comparison included a settings difference: Firmulate says Kimi K3 ran without an effort parameter and used the API default, while the other models ran at xhigh. That caveat is relevant when interpreting the standings. The league records behavior in one designed scenario, with a particular setup and scoring system; it does not by itself settle how the models would perform across other companies or under different conditions.
““Same diagnosis, same pitch — no signature.””
— Firmulate
Limits of the Published Results
The reported rankings come from one simulated company and one difficult week. The supplied account does not detail the full scoring rubric, the precise actions behind every score, or how results might change with different scenarios, data or model settings. The effort-parameter difference also complicates direct comparison between Kimi K3 and the other participants.
Firmulate says its enterprise pilot will use read-only exports and will not write back to real systems. The available description does not specify which data fields a pilot requires, how data would be protected or retained, how scenarios would be selected, or what evaluation standards would govern a board report. No pilot results are provided, so the reported simulation findings do not establish that a model is ready to perform business tasks autonomously.
Company-Specific Pilots Ahead
Firmulate is inviting companies to discuss pilots built around a read-only export of their business data. The proposed next step is to run crisis scenarios and deliver a board report covering model rankings and weaknesses in the company’s playbooks. The company says those tests would not write to live systems.
Readers can follow the synthetic company at firmulate.com/live and review the full results at firmulate.com/benchmarks.html. Firmulate lists its pilot page and contact@firmulate.com for companies interested in discussing a trial. Whether the approach identifies failures that transfer to day-to-day operations remains to be shown through company-specific evaluations.
Source: ThorstenMeyerAI.com
Key Questions
What did the Firmulate test measure?
It measured how five AI models handled a simulated software company during a difficult week, including crises, a sales opportunity, staged manipulation attempts and a restricted action.
Which model ranked highest?
Firmulate’s final standings put gpt-5.6-sol at 95, followed by Kimi K3 at 93. The scores apply to this experiment, and Kimi K3 used the API’s default effort setting while the others ran at xhigh.
Did all the models close the €55,000 deal?
No. Firmulate says only two signed the deal, despite all five identifying the crises. The key competitive information was buried in the company’s files.
How is the proposed enterprise pilot designed?
Firmulate says the pilot would run scenarios against a read-only export of a company’s data and produce a board report. It says the pilot would not write back to real systems.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
