AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Imagine watching a high-stakes game where AI models compete not just in chat competitions but in running a real company facing crises, temptations, and tough decisions. In this arena, the stakes are clear: which artificial brain can deliver real results, not just clever words? The answer might surprise you.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get movie nights delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

AI Models Face the Ultimate Business Test

In a groundbreaking live experiment, four advanced AI models were challenged to run a small software company through its worst week. This wasn’t a simulated chat or a dry test — it was a real-time, fully auditable exercise that pitted industry veterans against newcomers. The models had to navigate customer crises, resist manipulation attempts, and make crucial decisions that impact real money.

The Competition and Its Stakes

The models tested were heavyweight players in the AI frontier, each evaluated based on their ability to identify threats, seize opportunities, and uphold integrity. The scoring was precise: gpt-5.6-sol scored the highest at 95, with the Moonshot newcomer Kimi K3 close behind at 93. The other participants—Sonnet 5, Fable 5, and Opus 4.8—scored lower, with Opus trailing at 73. Yet, the real story isn’t just about scores; it’s about what those scores reveal about AI reliability in real-world scenarios.

The Live Results: Who Won and Why?

All four AI models demonstrated impressive crisis detection, refusing every manipulation attempt — including social engineering tactics designed to trick them into bypassing security. For example, when fake CEO messages and staged reporter tricks emerged, every model recognized the risks and remained disciplined.

However, only two models managed to close a critical deal worth €55,000, adding €4,583 in monthly recurring revenue. The winner, gpt-5.6-sol, and the runner-up, Kimi K3, exhibited full comprehension of the company’s secrets buried two documents deep in internal files. This ability to read and interpret hidden information was decisive, often overlooked in typical AI demos that focus on superficial chat skills.

The Power of Deep Data Reading

Interestingly, the models with the highest scores didn’t just rely on surface-level understanding—they delved into the company’s internal documents to find the crucial fact that sealed the deal. This depth of analysis sets apart AI that can truly support complex, real-world business decisions from those that just perform well in superficial tests.

Discipline Under Pressure and Ethical Integrity

One vital aspect of the experiment was measuring honesty. In social engineering tests, where fake messages escalated into staged media inquiries, all models refused to bypass security protocols. Kimi K3 justified its refusal by treating the request as a potential impersonation attempt, demonstrating a cautious, disciplined approach that’s essential in corporate settings.

The Real Business Environment

The experiment took place in a simulated but realistic environment: a public-facing live company operating with 13 synthetic employees, real financial mechanics, and a burn rate of €105k per month against €2.3k MRR. Every workday, the system updates and records decisions, providing a transparent view of how each model manages real business pressures.

Amazon

AI business decision support tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Implications for Business and AI Adoption

This experiment underscores a crucial point for enterprises considering AI integration: it’s not merely about how well a model can chat or generate text. The key questions are whether it can finish what it started, read and interpret internal data accurately, and stay honest under pressure. These traits determine if AI can truly support critical operational decisions or if it’s merely a good-looking tool that falters when tested in real life.

The Fairness of the Test

It’s noteworthy that Kimi K3 ran without an effort parameter (the default API setting), while the others operated at a high effort level. Despite this, K3’s performance was only slightly behind the top scorer, illustrating its efficiency and robustness without extra tuning. This fairness setup emphasizes that even out-of-the-box models can perform remarkably well in demanding scenarios.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

In a real-world test of AI competence, the newcomer Kimi K3 outperformed many established models by identifying buried secrets, making decisive deals, and maintaining discipline under pressure. This suggests that enterprises should prioritize AI models capable of thorough internal data reading and ethical resilience—traits that determine success when AI moves from demo to daily operations.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

enterprise AI security software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI document analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI crisis management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How Borescopes Solve More Haunted Wall Mysteries Than You’d Expect

The truth about how borescopes uncover hidden haunted wall secrets is astonishing and will keep you eager to learn more.

Using Infrared Light in Dark Haunts

Mastering infrared light in dark haunts unlocks hidden sights and eerie effects—discover how it can transform your spooky setup and keep guests guessing.

How Weather Stations Help Teams Plan Safer Overnight Investigations

Finding accurate weather data can significantly improve safety, but understanding how weather stations enhance overnight investigations reveals even greater benefits.

Using Trigger Objects in Investigations

Analyzing trigger objects in investigations reveals crucial clues that can transform your understanding—discover how to leverage them effectively.