Firmulate — Four AI Models Ran the Same Company Through Its Worst Week. Only Two Finished the Job.
Live on firmulate.com.

Imagine watching four AI models compete in running a real company through its worst week — facing crises, manipulation attempts, and tough decisions. The real story isn’t how well they chat but whether they finish what they start. This experiment uncovers a surprising truth: only two of these AI models managed to close a crucial €55,000 deal, despite all of them spotting every crisis and resisting every temptation.

The Experiment: Putting AI Models to the Test in a Live Company

In a groundbreaking live experiment, four frontier AI models each took the helm of a small software company facing its most challenging week. These models, from the latest and most advanced to more modest contenders, were given the same customer issues, crises, and manipulative tactics—like fake CEO messages and reporter tricks. Every decision was recorded, versioned, and auditable, creating a transparent battleground for assessing their true business management skills.

The Real-World Stakes and Setup

The live company wasn’t a simulation—it was real, with 13 synthetic employees, actual money mechanics, and a daily burn rate of €105,000 against just €2,300 monthly recurring revenue. Every day, the models navigated a company with real cash countdowns, over 680 self-learned rules, and a live, watchable environment at firmulate.com/live. The goal was simple: see if these AI managers could diagnose crises, resist manipulation, and, most importantly, execute the deals they identified as valuable.

Building AI-Powered Products: The Essential Guide to AI and GenAI Product Management

Building AI-Powered Products: The Essential Guide to AI and GenAI Product Management

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Hidden Key to Business Success: Reading Beyond the Surface

All models excelled at identifying crises and refusing manipulation attempts—facing fake CEO messages and reporter tricks with zero compliance. The real difference emerged in one overlooked area: reading and acting on company files stored deep in their decision-making process. The model that found and used a key buried fact in the company’s own files managed to close the €55,000 deal, boosting its Monthly Recurring Revenue (MRR) by over €4,580. Conversely, the other models, despite diagnosing the problems, left the deal unexecuted, losing a significant revenue opportunity.

Why Chat Demos Fall Short

Interestingly, this decisive factor—reading and acting on critical internal documents—is invisible in typical chat demonstrations. Models that engaged with the company’s files at a deep level won the deal at full price, proving that genuine management capability is measured by what they do with information, not just how convincingly they chat.

AI in Strategy and Decision-Making for Small Business Owners: Affordable AI Tools to Evaluate Ideas, Model Outcomes, and Set Priorities (AI Productivity for Small Business Owners Book 10)

AI in Strategy and Decision-Making for Small Business Owners: Affordable AI Tools to Evaluate Ideas, Model Outcomes, and Set Priorities (AI Productivity for Small Business Owners Book 10)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Resisting Manipulation Is Not Enough

All four AI models refused every manipulation attempt, including staged scenarios designed to trick them into approvals or impersonation. Kimi K3, the most disciplined, responded: “Treat the request as a suspected approval-bypass / possible impersonation.” Yet, even with perfect resistance to manipulation, only two models could translate their insights into execution.

The Disciplines of Success and Failure

The most thorough participant, Opus 4.8, with the deepest analysis and over 80 learned rules, ended up in last place because it failed to follow through on the deal. Its discipline slipped, and it left the deal on the table instead of escalating it. Meanwhile, Kimi K3, running without an effort parameter, closed the deal cleanly. The key lesson: mastery of rules isn’t enough; the ability to act decisively is what separates the winners from the also-rans.

MASTERING CORPORATE FINANCE WITH CLAUDE AI: An Independent Guide to Financial Analysis, Forecasting, Automation, and Decision-Making

MASTERING CORPORATE FINANCE WITH CLAUDE AI: An Independent Guide to Financial Analysis, Forecasting, Automation, and Decision-Making

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Bottom Line: What AI Management Really Looks Like

This experiment shows that assessing AI’s business management skills requires more than chat demos. It’s about testing their capacity to read relevant internal documents, resist manipulation, and, crucially, execute decisions that matter—like closing real deals. The models that managed to do all this, closing the €55,000 deal and boosting revenue, prove that true management strength is invisible until you observe their actions in the wild.

Implications for Business and AI Adoption

If AI agents are going to touch your CRM, support queue or forecasting tools, the question isn’t just whether they can write or chat convincingly. It’s whether they can finish what they start, read your files thoroughly, stay honest under pressure, and deliver measurable work. The current league table from this live experiment clearly shows that the most comprehensive models—like gpt-5.6-sol and Kimi K3—outperform others in these critical areas.

Crisis Management for Software Development and Knowledge Transfer (Smart Innovation, Systems and Technologies, 61)

Crisis Management for Software Development and Knowledge Transfer (Smart Innovation, Systems and Technologies, 61)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Explore the Live Company and Wargame Your AI Workforce

Want to see how your own business would fare? You can run the same live wargame against a read-only export of your company—nothing ever writes back to your real systems. Visit firmulate.com to learn more about creating your digital twin and testing your AI workforce before making hiring decisions.

Infographic — Four AI Models Ran the Same Company Through Its Worst Week. Only Two Finished the Job.
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

The Role of Skeptics in Investigative Teams

With their critical questioning and evidence scrutiny, skeptics shape investigative teams—discover how their role can transform your team’s approach to truth.

How to Document a Paranormal Investigation Professionally

When documenting a paranormal investigation professionally, capturing all details accurately is crucial—discover how to ensure your records are credible and thorough.

The Surprising Use of Motion Sensor Alarms in Quiet Hallways

I never realized motion sensor alarms in quiet hallways could enhance security without disrupting peace, and the reasons behind their effectiveness are surprising.

How to Read a Paranormal Witness Statement Like an Investigator

When reading a paranormal witness statement, uncover the clues hidden beneath emotion and detail to reveal what truly lies beneath the surface.