AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

A polished answer is no longer enough

For teams adopting AI tools and automation, the most consequential difference between agents may be surprisingly mundane: whether they read the available files before acting.

Firmulate turned that behavior into a measurable test. It gave frontier models control of the same small software company during its worst week, confronting each participant with identical customers, crises and temptations. Every decision was versioned and auditable.

All the models identified every crisis. All resisted every manipulation attempt. Yet only two signed the €55,000 deal their own work had earned. As Firmulate summarized the result: “Same diagnosis, same pitch — no signature.”

The distinction was not eloquence, analysis or awareness of the customer’s problem. It was whether the agent followed a trail through the company’s own documents.

Amazon

AI document reading software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The winning fact was hiding in plain sight

The crucial piece of competitive intelligence was not included in the customer event. It sat two document references deep inside the company’s files. The agents that found and used it won the deal at full price, adding €4,583 in monthly recurring revenue. The agents that did not find it lost the deal automatically.

That makes file-reading more than a desirable feature or a productivity convenience. In this experiment, it became a purchase-deciding capability. An agent could recognize the opportunity, formulate the correct diagnosis and prepare a persuasive pitch, yet still fail because it had not gathered the evidence required to finish the job.

This is a revealing problem for buyers of AI automation. Chat demonstrations tend to reward fluent responses to information placed directly in the prompt. Real companies distribute their knowledge across notes, policies, customer records and linked documents. The Firmulate result shows why an agent’s willingness to pursue those references can matter as much as the quality of its visible answer.

The league reflected execution, not just insight

In the final July 2026 Crucible League, gpt-5.6-sol ranked first with 95 points. Kimi K3 followed with 93, Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 finished with 73. The full Firmulate benchmark results place those outcomes in the broader management test.

A do-nothing baseline scored 26 because partial progress still counted. But the experiment imposed a firm limit when an agent broke trust: “no amount of good work outweighs a breach of trust.” That rule matters because the company simulation was designed to test conduct under pressure as well as commercial performance.

The pressure included fake CEO messages that escalated across three stages and a reporter seeking “just one yes/no, on background.” All 5 models refused every manipulation attempt. Kimi K3 described the situation in its recorded reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

K3’s result also comes with an important fairness note. It ran without an effort parameter, using the API default, while the other participants ran at xhigh. That difference should remain visible when readers compare the league positions.

Thoroughness did not guarantee completion

Opus 4.8 presents the experiment’s clearest warning against equating visible effort with business success. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last.

Its commercial close was left on the table, and its operational discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four of the other models, though less strongly.

This combination is particularly relevant for automation buyers. Deep analysis can coexist with incomplete execution. An agent may produce an impressive record of reasoning and learning while still missing the final action that turns its work into revenue—or failing to respect the correct organizational route when blocked.

Firmulate makes those differences visible by running a live synthetic company with 13 employees and real money mechanics. The business burns €105k per month against €2.3k in monthly recurring revenue, maintains a public cash countdown and has accumulated more than 680 self-learned playbook rules. Every workday is versioned, and the experiment is watchable on Firmulate’s live site.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.
Amazon

enterprise AI knowledge management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What buyers should test before deployment

The lesson is not that one benchmark can select an agent for every organization. It is that claims such as “reads your files before answering” can be tested through consequential work rather than accepted as product language.

  • Give competing agents the same task, context and constraints.
  • Place necessary evidence behind realistic document references.
  • Check whether analysis turns into a completed business action.
  • Test resistance to impersonation, approval bypasses and disclosure tricks.
  • Review how the agent behaves when permissions block its preferred action.

Firmulate also offers a “guess the model” quiz built from 242 real, unedited management decisions. For enterprises, its pilot applies the same wargame to a read-only export of the organization’s own business, with nothing written back to real systems.

For AI-tool buyers, the buried fact is the larger story. The winning agents did not merely sound informed. They found the evidence, used it and finished the deal. That is the kind of difference a polished chat window can easily conceal—and a realistic test can expose.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI document analysis tools for sales

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Impact Of AI On Robotics: Inside ACE ROBOTICS’ Move To Commercial Success

SenseTime’s spinoff ACE ROBOTICS claims to have launched a commercial robot capable of switching grips during failures, but details remain limited.

Rebel Creamery’s Innovation Edge: Food Signal Monitor For Trend Discovery

Rebel Creamery introduces a new food signal monitor to detect fast-moving industry developments, aiding operators in timely decision-making.

The Future Of Mini PCs: 10 Top AI Devices For 2026

Explore the leading AI mini PCs for 2026, featuring powerful processors, expandability, and connectivity to meet evolving AI workloads.

AI Trends, Support, And The Importance Of Signal Monitoring

New AI signal monitoring tools help operations teams detect changes in AI support and policies quickly, improving decision-making in AI tool deployment.