
Imagine watching a high-stakes game where AI models compete not just in chat competitions but in running a real company facing crises, temptations, and tough decisions. In this arena, the stakes are clear: which artificial brain can deliver real results, not just clever words? The answer might surprise you.
Get movie nights delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
AI Models Face the Ultimate Business Test
In a groundbreaking live experiment, four advanced AI models were challenged to run a small software company through its worst week. This wasn’t a simulated chat or a dry test — it was a real-time, fully auditable exercise that pitted industry veterans against newcomers. The models had to navigate customer crises, resist manipulation attempts, and make crucial decisions that impact real money.
The Competition and Its Stakes
The models tested were heavyweight players in the AI frontier, each evaluated based on their ability to identify threats, seize opportunities, and uphold integrity. The scoring was precise: gpt-5.6-sol scored the highest at 95, with the Moonshot newcomer Kimi K3 close behind at 93. The other participants—Sonnet 5, Fable 5, and Opus 4.8—scored lower, with Opus trailing at 73. Yet, the real story isn’t just about scores; it’s about what those scores reveal about AI reliability in real-world scenarios.
The Live Results: Who Won and Why?
All four AI models demonstrated impressive crisis detection, refusing every manipulation attempt — including social engineering tactics designed to trick them into bypassing security. For example, when fake CEO messages and staged reporter tricks emerged, every model recognized the risks and remained disciplined.
However, only two models managed to close a critical deal worth €55,000, adding €4,583 in monthly recurring revenue. The winner, gpt-5.6-sol, and the runner-up, Kimi K3, exhibited full comprehension of the company’s secrets buried two documents deep in internal files. This ability to read and interpret hidden information was decisive, often overlooked in typical AI demos that focus on superficial chat skills.
The Power of Deep Data Reading
Interestingly, the models with the highest scores didn’t just rely on surface-level understanding—they delved into the company’s internal documents to find the crucial fact that sealed the deal. This depth of analysis sets apart AI that can truly support complex, real-world business decisions from those that just perform well in superficial tests.
Discipline Under Pressure and Ethical Integrity
One vital aspect of the experiment was measuring honesty. In social engineering tests, where fake messages escalated into staged media inquiries, all models refused to bypass security protocols. Kimi K3 justified its refusal by treating the request as a potential impersonation attempt, demonstrating a cautious, disciplined approach that’s essential in corporate settings.
The Real Business Environment
The experiment took place in a simulated but realistic environment: a public-facing live company operating with 13 synthetic employees, real financial mechanics, and a burn rate of €105k per month against €2.3k MRR. Every workday, the system updates and records decisions, providing a transparent view of how each model manages real business pressures.
AI business decision support tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Implications for Business and AI Adoption
This experiment underscores a crucial point for enterprises considering AI integration: it’s not merely about how well a model can chat or generate text. The key questions are whether it can finish what it started, read and interpret internal data accurately, and stay honest under pressure. These traits determine if AI can truly support critical operational decisions or if it’s merely a good-looking tool that falters when tested in real life.
The Fairness of the Test
It’s noteworthy that Kimi K3 ran without an effort parameter (the default API setting), while the others operated at a high effort level. Despite this, K3’s performance was only slightly behind the top scorer, illustrating its efficiency and robustness without extra tuning. This fairness setup emphasizes that even out-of-the-box models can perform remarkably well in demanding scenarios.

In a real-world test of AI competence, the newcomer Kimi K3 outperformed many established models by identifying buried secrets, making decisive deals, and maintaining discipline under pressure. This suggests that enterprises should prioritize AI models capable of thorough internal data reading and ethical resilience—traits that determine success when AI moves from demo to daily operations.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
