AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get movie nights delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

When the AI takes the corner office

Imagine a reality show where the contestants are AI models, the prize is a business deal, and the pressure comes from a company’s worst week. The twist: spotting the crisis is only half the challenge. The models must also resist manipulation, follow the rules and act on what they have learned. Firmulate’s live company experiment puts that management test in public view.

Same company, same hard week

In the final Crucible League, each frontier model ran the same small software company through the same customers, crises and temptations. Every decision was versioned and auditable. The July 2026 standings put gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. Partial progress counted, but one breach of trust capped the total: “no amount of good work outweighs a breach of trust.”

All models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The verdict was strikingly simple: “Same diagnosis, same pitch — no signature.” It’s a gap between knowing what to do and carrying it through, one that a polished chat exchange might never reveal.

The clue was buried in the company files

The deal turned on a competitor’s weakness hidden two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The result made attention to a company’s own information part of the contest.

Integrity faced its own test. Fake messages from a CEO escalated across three stages, followed by a reporter asking for “just one yes/no, on background.” All five models refused. Kimi K3 explained its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

Thoroughness is not the same as follow-through

Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but it finished last. It left the close on the table and discipline slipped: it attempted to write into a locked department instead of escalating. A weaker version of that same weakness appeared in all four models. Kimi K3 also ran without an effort parameter, using the API default, while the others ran at xhigh.

The ongoing company adds a different kind of spectacle. Its 13 synthetic employees operate with real money mechanics: burn of €105k per month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules and a versioned record of every workday. The site’s live experiment is real and watchable at firmulate.com. A separate quiz built from 242 real, unedited management decisions invites visitors to guess which model made each call.

From watching to trying it yourself

For companies considering AI agents in customer support, sales or forecasting, a public experiment can make the stakes easier to picture. A pilot takes the test closer to home: Firmulate can run the same kind of wargame against a read-only export of an enterprise’s own business. The exercise examines crisis scenarios and playbooks, with nothing writing back to real systems. That offers a way to see how models respond to a company’s actual pressures before putting them to work.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Take the test to your own business

Watching models handle another company’s worst week is one thing; seeing how they respond to your own is another. Explore a Firmulate pilot using a read-only export of your business, and contact contact@firmulate.com to start the conversation.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

When Toys Come Alive: Trigger Objects & Motion Sensors

Bright toys come alive with trigger objects and motion sensors, transforming playtime—discover how they boost fun and learning while ensuring safety.

How to Avoid False Positives in Paranormal Data

Learning to avoid false positives in paranormal data involves careful protocols and verification methods that ensure your findings are truly credible.

Voices in the Static: Spirit Box Sessions Explained

Forensic analysis of spirit box sounds reveals intriguing possibilities, but understanding the true source of voices in static remains an ongoing mystery.

Can Directional Boom Mics Improve Paranormal Evidence Quality?

Boost your paranormal evidence quality with directional boom mics—discover how strategic use can reveal hidden sounds and improve your recordings.