AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine a world where AI managers are tested not just on how well they chat, but on their ability to run a business under pressure—honest, disciplined, reliable. Sounds like science fiction? It’s actually happening, and it’s changing how we judge AI’s true worth.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get movie nights delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The New Standard in AI Evaluation

Enter the latest experiment by Firmulate, a pioneering public benchmark designed to realistically assess AI models as management leaders. Unlike conventional tests that focus on fluent speech or clever tricks, this benchmark simulates a small software company’s worst week—full of crises, manipulative temptations, and high-stakes decisions. Every model faces identical scenarios: demanding customers, internal crises, and attempts to bend the rules.

Beyond the Surface: What the Scores Really Say

At first glance, the scores are impressive: the top model, GPT-5.6, scores 95 out of 100, while others like Kimi K3 and Sonnet 5 follow close behind. Yet, what truly matters is the baseline: a do-nothing approach scores 26 points. Why isn’t that zero? Because even doing the bare minimum—like reading documents or refusing manipulative requests—earns partial credit. This approach recognizes that in real-world management, partial progress and consistent honesty are crucial.

The Importance of Trust and Discipline

One key rule in the benchmark is that even a single breach of trust caps the total score. For example, models that read a file containing a critical fact and leverage it to close a lucrative deal end up with high scores—up to full credit. Meanwhile, models that miss this buried fact during their analysis or slip up in discipline leave money on the table, illustrating a clear link between honesty, thoroughness, and business success.

Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Honesty Under Pressure — The Critical Test

The experiment isn’t just about intelligence; it’s about integrity. All four models spotted every crisis and refused manipulative tricks like fake CEO messages or reporter tricks. For instance, when faced with staged social engineering efforts, each AI declined to escalate or approve suspicious requests, adhering to protocols like “treat the request as a suspected impersonation.” This consistent refusal underscores an essential trait: trustworthiness.

Real Business Mechanics, Not Just Words

The live simulation runs a small enterprise with 13 synthetic employees, handling real money mechanics—burning €105k a month against a revenue of just €2,300. The process involves over 680 self-learned rules and versioning every decision. Watching the experiment unfold at firmulate.com/live reveals AI’s capacity (or lack thereof) to uphold discipline amidst financial pressure.

Amazon

business decision making AI tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Surprising Weaknesses and What They Reveal

Despite all being disciplined enough to refuse manipulation, models like Opus 4.8 showed vulnerabilities. Even the most thorough participant—who learned over 80 rules and performed the deepest analyses—failed to close a deal, leaving potential revenue unclaimed. This highlights a critical insight: thoroughness alone doesn’t guarantee success; execution and discipline matter.

Fairness and Variations in Setup

The experiment also reveals how setups influence outcomes. For example, K3 ran without an effort parameter, which might have affected its performance, compared to others running at high effort levels. This provides context on how different configurations impact decision-making and results.

Amazon

AI ethics and trustworthiness assessment

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why This Matters for Your Business

Ultimately, this benchmark emphasizes a vital point: when deploying AI in real-world roles—such as managing CRM, support, or forecasting—the question isn’t just about linguistic prowess. It’s whether the AI can finish what it starts, read critical files, and stay honest under pressure. Partial progress and discipline aren’t optional—they’re essential for trustworthy management.

Join the Exploration

For enterprise leaders curious about testing their own AI tools, Firmulate offers a platform to run the same simulated week against their systems—without any risk to actual operations. This approach lets companies see how their AI workforce performs in realistic, high-pressure scenarios, helping them choose solutions that are truly reliable.

Amazon

AI management decision benchmarks

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Bottom Line: A More Honest Benchmark

This experiment exposes a simple truth: trustworthiness and thoroughness are as vital as technical sophistication. The benchmark’s modest baseline score of 26 points reminds us that AI’s real value lies not in clever words, but in consistent, honest, and disciplined performance—especially when the stakes are high.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

When Toys Come Alive: Trigger Objects & Motion Sensors

Bright toys come alive with trigger objects and motion sensors, transforming playtime—discover how they boost fun and learning while ensuring safety.

EVP Session Guide: Capturing Spirit Voices

Join us as we uncover expert tips to enhance your EVP sessions and finally hear what spirits want to say.

How Weather Stations Help Teams Plan Safer Overnight Investigations

Finding accurate weather data can significantly improve safety, but understanding how weather stations enhance overnight investigations reveals even greater benefits.

How Data Logging EMF Detectors Help Track Repeating Events

Discover how data logging EMF detectors track repeating events over time, revealing patterns that can help you better understand and manage electromagnetic interference.