AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.

What happens when automation has to run the business?

Most AI tools are demonstrated in forgiving conditions: write an email, summarize a meeting, produce a plan. Firmulate offers readers interested in AI tools and automation a harsher test. Its software company must handle customers, crises and financial pressure while the public watches.

The live experiment has 13 synthetic employees and real money mechanics. It burns €105k a month against €2.3k in monthly recurring revenue, displays a public cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned. The result is build-in-public pushed beyond product updates and revenue charts: the company’s struggle to survive becomes the product, the test and the running story.

AI in Property Management: A Practical, Unboring Look at Artificial Intelligence in the Multifamily Industry

AI in Property Management: A Practical, Unboring Look at Artificial Intelligence in the Multifamily Industry

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A business under pressure, not a chatbot on display

Firmulate’s Crucible League put frontier models in charge of the same small software company during its worst week. Each participant encountered the same customers, crises and temptations. Every decision was versioned and auditable, making it possible to compare management behavior rather than polished conversational output.

The final July 2026 standings placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. One boundary remained absolute: a single breach of trust capped the total because “no amount of good work outweighs a breach of trust.”

The reassuring finding was that all models identified every crisis and rejected every manipulation attempt. The more revealing result was that only two signed the €55,000 deal their own work had earned. Firmulate’s summary captures the gap: “Same diagnosis, same pitch — no signature.”

The decisive fact was buried in ordinary company material

The deal did not turn on eloquence alone. A decisive weakness in the competitor’s position was hidden two document references deep in the company’s own files rather than presented in the customer event. The models that found and used it won the deal at full price, adding €4,583 in monthly recurring revenue.

That distinction should matter to anyone evaluating AI automation. Spotting an incoming issue is not the same as completing the surrounding work. A useful business agent must consult the available record, connect information across documents and carry a justified recommendation through to an outcome. In Firmulate’s experiment, models could reach the same diagnosis and even produce the same pitch, yet still diverge at the moment when action mattered.

Manipulation met a consistent refusal

The week also included fake CEO messages escalating over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”

This is significant because the pressure resembled recognizable workplace tactics rather than an abstract safety puzzle. Authority was invoked, urgency escalated and confidentiality was framed as informal. The field held the line across every attempt, showing that strong completion and resistance to manipulation can be evaluated in the same business setting.

Thoroughness did not guarantee the best result

Opus 4.8 was the most thorough participant. It produced the deepest analyses and added 80 learned rules, yet finished last in the league. The close was left on the table, while discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four of the others, although less strongly.

That profile complicates a common assumption about AI work: more analysis does not automatically produce better management. Thoroughness can be valuable, but the experiment’s outcome depended on disciplined follow-through, correct escalation and finishing the commercially important task.

There is also an important fairness note. Kimi K3 ran with the API default because it had no effort parameter, while the other models ran at xhigh. Its second-place score should therefore be read with that difference in mind.

A public company creates daily evidence

Firmulate’s live company turns these questions into an ongoing operating record. Its financial imbalance is visible, its cash countdown is public and its workdays continue to generate material. Visitors can also read what the synthetic employees say, adding a human-readable layer to the versioned decisions.

The broader portrait is unusual: a company with no human employees, a steep gap between burn and recurring revenue, and a growing body of rules learned from its own work. It is not presented as a fictional simulation. It is real, watchable software using business pressure to expose where AI management succeeds, stalls or loses discipline.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.
Amazon

business AI decision support tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The lesson for automation buyers

Firmulate’s public experiment suggests that the meaningful test of an AI worker begins after it produces a plausible answer. Can it find a critical fact buried in company material? Can it resist an apparent executive’s attempt to bypass approval? Can it respect boundaries, escalate correctly and complete the deal?

For readers considering AI tools, that is the practical value of this live-company story. The leaderboard is interesting, but the daily record is more consequential: automation can recognize the right move without making it, and the distance between insight and completion can be worth a €55,000 contract.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


AI for Public Relations: A How-To Guide for Implementation and Management

AI for Public Relations: A How-To Guide for Implementation and Management

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Driver Safety Innovation: Aftermarket Fatigue Detection Devices

New phone-based fatigue detection app aims to warn long-commute drivers of drowsiness, offering a potential safety upgrade for older vehicles without built-in tech.

Mark Zuckerberg Tells Staff That AI Agents Haven’t Progressed Enough

Facebook CEO Mark Zuckerberg told staff that AI agents are not sufficiently developed, signaling cautious outlook on AI progress.

The bridge. Why the AI buildout runs on a nuclear story and a gas reality.

Analysis of the AI industry’s energy strategy reveals a nuclear procurement rush contrasted by immediate reliance on gas for power needs, highlighting a timeline mismatch.

Cutrova: Edit the Words, Not the Timeline

Cutrova launches a local-first, transcript-based video editor that simplifies editing by focusing on words rather than timelines, emphasizing privacy and control.