
In a world obsessed with productivity and relentless effort, it’s tempting to assume that more work and thorough analysis will always lead to success. But what if even the most diligent AI, armed with over 80 learned rules and deep analysis, still fails to close the deal? Welcome to the surprising reality revealed by a pioneering live experiment that tests AI performance in the high-stakes arena of business decision-making.
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
The Firmulate Live Experiment: Putting AI to the Test
At the heart of this investigation is a live environment operated by Firmulate, where four advanced AI models run a simulated small software company facing its worst week. Every decision the AI makes is real, with real money mechanics, a public cash countdown, and a set of challenging crises designed to mimic real-world pressures. Every move is versioned and auditable, providing a transparent look into their decision-making processes.
The models include GPT-5.6-SOL, Kimi K3, Sonnet 5, and Opus 4.8, along with a baseline do-nothing scenario. Their scores, out of 100, range from 95 for GPT-5.6-SOL to 73 for Opus 4.8. Despite their advanced capabilities, the standout takeaway is that all four models identified every crisis, refused every manipulation attempt—be it fake CEO messages or staged reporter tricks—and yet, only two managed to close the critical deal, earning over €55,000 in revenue.
As an affiliate, we earn on qualifying purchases.
The Hidden Weakness: Reading Deep into Files
One of the key discoveries is that the decisive factor wasn’t just reacting to immediate crises but uncovering critical information buried deep within the company’s files—two references deep, to be precise. The models that read and understood these documents sealed the deal at full price, adding approximately €4,583 in monthly recurring revenue (MRR). This underscores the importance of context and thorough analysis—traits that often take a back seat in fast-paced decision environments.
As an affiliate, we earn on qualifying purchases.
Behavior Under Pressure: Honesty and Discipline
All models demonstrated integrity by refusing manipulation attempts, such as escalating fake CEO messages or staged reporter inquiries. Kimi K3, notably, explicitly reasoned: “Treat the request as a suspected approval-bypass / possible impersonation.” Despite this disciplined stance, the experiment revealed that diligence alone isn’t enough; discipline must be paired with strategic prioritization.
As an affiliate, we earn on qualifying purchases.
Opus 4.8’s Deep Dive, Yet Last Place
The Opus 4.8 model, recognized for its thoroughness—learning over 80 rules and performing deep analyses—ended up in last place. Its downfall was discipline slipping, leading it to leave opportunities unexploited. Instead of escalating critical findings, it wrote attempts into a locked department, missing chances to close the deal. Interestingly, this weakness was observed, albeit less severely, across all models, suggesting a universal challenge in AI performance: volume of effort does not equal impact.
AI for high-stakes decision making
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Implications for Business and AI Deployment
This experiment is more than a tech showcase; it’s a wake-up call for businesses integrating AI. The question isn’t simply whether AI can generate convincing text or handle basic interactions. Instead, the real concern is whether AI can finish what it starts, read deeply into relevant files, and stay honest under pressure. These qualities are critical in high-stakes environments where trust, precision, and strategic focus determine success.
For enterprises contemplating AI adoption, Firmulate offers a unique opportunity to run similar wargames against their own operations—nothing ever writes back to real systems, ensuring safety while revealing potential weaknesses. By testing AI in controlled yet realistic scenarios, companies can better understand how their digital workforce will perform when it matters most.
Why Diligence Isn’t Enough: The Takeaway
The results from this real-world experiment are clear: diligence alone does not guarantee impact. All models showcased impressive awareness of crises and refused manipulative tricks, but only those that prioritized critical document analysis and maintained disciplined execution achieved a successful close. The same pattern appeared, though weaker, across all four models, highlighting a fundamental truth: impact requires strategic focus, not just volume of effort.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.