
Imagine watching a high-stakes game where competitors don’t just score points— they also navigate crises, dodge manipulations, and stay honest under pressure. Now, picture AI agents playing this game in real time, managing a company’s worst week. That’s exactly what the latest experiment from Firmulate is doing, and it’s revealing a hidden truth: the real test of AI isn’t how well it chats, but how well it manages chaos.
Stepping into the Business Arena
In an era where AI chatbots are judged mainly by their ability to generate convincing language, a new frontier is emerging—measuring management quality. The goal isn’t just to produce accurate answers but to see if these digital agents can handle real-world crises, make ethical decisions under pressure, and stick to their analysis—even when temptations to cheat or manipulate arise.
AI management decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Experiment: Putting AI to the Test
The firmulate.com live experiment set four advanced AI models against a simulated small software company facing its worst week. Every model was given identical challenges: tough customer crises, internal manipulations, and schemes to bypass controls. Each decision was recorded and auditable, ensuring transparency and fairness.
business crisis simulation AI tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Results: A Clear Gap in Management Skills
All four AI models demonstrated impressive capabilities. They identified every crisis and refused every attempt at manipulation, showing high integrity. However, only two teams managed to close the deal, earning €55,000 in revenue. The other two failed to finalize the agreement, despite making the same diagnoses and pitches. This gap was not caused by poor decision-making but by deeper issues: reading and understanding company documents, following protocols, and resisting shortcuts.
ethical AI decision support systems
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Hidden Weakness
Further analysis revealed that the decisive advantage came from reading two document references deep within the company’s files. Models that accessed and understood these references secured the deal at full price—more than €4,500 monthly recurring revenue (MRR). Ironically, the models that skipped this step lost the opportunity, emphasizing that surface-level chat prowess isn’t enough to succeed in complex management tasks.
As an affiliate, we earn on qualifying purchases.
Handling Social Engineering and Ethical Dilemmas
Another test involved fake CEO messages escalating through multiple stages and a reporter trick asking for a discreet approval. All five models refused to cooperate, treating these as suspicious attempts at impersonation or bypassing approval processes. Kimi K3’s explanation was clear: “Treat the request as a suspected approval-bypass / possible impersonation.” This shows the models’ capacity for ethical judgment, not just answer generation.
The Live Business: A Real Company in Action
The experiment runs a real, money-losing software company with 13 synthetic employees, real cash mechanics, and over 680 self-learned rules. Every workday, the company’s decisions are versioned and observable—like a live laboratory of AI management in action. Visitors can watch this ongoing simulation at firmulate.com/live.
What This Means for Business
While current AI models excel in chat demos, their true test lies in management capabilities: finishing tasks, reading critical files, maintaining honesty, and handling pressure. The experiment underscores the importance of management quality over superficial answers—a crucial insight for businesses deploying AI in support, CRM, or forecasting roles.
Beyond the Scores: The Management Competency League
The experiment’s leaderboard shows GPT-5.6-sol leading with a score of 95, followed by the newcomer Kimi K3 at 93, then Sonnet 5 at 88, and another Sonnet at 77. The scores reflect each model’s ability to close deals, resist manipulation, and read company documents. Notably, the models ran at different effort levels, influencing their discipline and thoroughness.
The Future of AI in Business
As AI agents become more integrated into daily operations, understanding their management capabilities becomes vital. This experiment by Firmulate is a step toward evaluating AI not just by what they say but by what they do—under pressure, ethically, and strategically. For enterprises, the key question is: can your AI handle the messy, unpredictable realities of running a business?

AI’s true value lies in management quality — reading, decision-making, honesty, and resilience— not just chat prowess. The Firmulate live experiment reveals that sophisticated models can succeed in complex, pressure-filled scenarios, but only if they go beyond surface answers. Businesses should evaluate AI by its ability to finish what it starts, interpret critical documents, and stay honest under stress, shaping the future of trustworthy automation.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html