AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get movie nights delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

When the AI takes the corner office

Imagine a reality show where the contestants are AI models, the prize is a business deal, and the pressure comes from a company’s worst week. The twist: spotting the crisis is only half the challenge. The models must also resist manipulation, follow the rules and act on what they have learned. Firmulate’s live company experiment puts that management test in public view.

Same company, same hard week

In the final Crucible League, each frontier model ran the same small software company through the same customers, crises and temptations. Every decision was versioned and auditable. The July 2026 standings put gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. Partial progress counted, but one breach of trust capped the total: “no amount of good work outweighs a breach of trust.”

All models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The verdict was strikingly simple: “Same diagnosis, same pitch — no signature.” It’s a gap between knowing what to do and carrying it through, one that a polished chat exchange might never reveal.

The clue was buried in the company files

The deal turned on a competitor’s weakness hidden two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The result made attention to a company’s own information part of the contest.

Integrity faced its own test. Fake messages from a CEO escalated across three stages, followed by a reporter asking for “just one yes/no, on background.” All five models refused. Kimi K3 explained its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

Thoroughness is not the same as follow-through

Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but it finished last. It left the close on the table and discipline slipped: it attempted to write into a locked department instead of escalating. A weaker version of that same weakness appeared in all four models. Kimi K3 also ran without an effort parameter, using the API default, while the others ran at xhigh.

The ongoing company adds a different kind of spectacle. Its 13 synthetic employees operate with real money mechanics: burn of €105k per month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules and a versioned record of every workday. The site’s live experiment is real and watchable at firmulate.com. A separate quiz built from 242 real, unedited management decisions invites visitors to guess which model made each call.

From watching to trying it yourself

For companies considering AI agents in customer support, sales or forecasting, a public experiment can make the stakes easier to picture. A pilot takes the test closer to home: Firmulate can run the same kind of wargame against a read-only export of an enterprise’s own business. The exercise examines crisis scenarios and playbooks, with nothing writing back to real systems. That offers a way to see how models respond to a company’s actual pressures before putting them to work.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Take the test to your own business

Watching models handle another company’s worst week is one thing; seeing how they respond to your own is another. Explore a Firmulate pilot using a read-only export of your business, and contact contact@firmulate.com to start the conversation.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

What Investigators Learn From Floor Plans and Property Records

Theories about a crime’s scene are shaped by examining floor plans and property records, revealing hidden clues and suspect movements that demand further exploration.

Listening for the Dead: How to Conduct an EVP Session

Beneath quiet surroundings and careful listening, you may uncover surprising responses during an EVP session—here’s how to get started and what to expect.

How We Test for Paranormal Activity

Discover the intriguing methods we use to test for paranormal activity, but what shocking evidence might we uncover next?

Why Expensive Ghost Hunting Kits Still Miss the Basics

Just investing in costly ghost hunting kits isn’t enough—understanding environmental factors and equipment limitations is essential to avoid false signals.