firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine a test where the worst possible AI performs just slightly above doing nothing at all. Surprisingly, that’s exactly what happens in the latest business AI benchmark — and it reveals a lot about trust, progress, and how ready AI really is for your company.

For listenersOffer from Amazon

Turn the school run and nap time into listening time

  • Thousands of audiobooks, podcasts and originals
  • Listen on your phone, tablet or Echo — also offline
  • Cancel anytime
Try Audible free Free trial for new members
As an affiliate, we earn on qualifying purchases.

At first glance, it might seem strange that an AI model with no real effort can score as high as 26 points out of a possible 100 in a business management test. But this baseline score isn’t a fluke — it’s a deliberate part of the assessment design, giving us a clear picture of what “doing nothing” actually looks like in practice.

Why Does Doing Nothing Score 26?

The scoring system in this benchmark recognizes that even a passive or indifferent AI can avoid making mistakes or being manipulated. In other words, if an AI refuses to act or manipulate, it automatically earns some points for simply not making things worse. This baseline of 26 points sets a realistic floor: an AI that does nothing is better than one that actively causes trouble or breaches trust.

The Partial Progress Paradox

Another key lesson from this test is that partial progress counts. For example, models that read and understand a critical document deep inside the company files can close a deal worth €4,583 in monthly recurring revenue, even if they stumble in other areas. This means that an AI’s ability to find and act on crucial hidden information can be worth more than its performance in superficial tasks.

Trust Matters Too

Perhaps most revealing is that a single breach of trust caps the total score, regardless of other achievements. If an AI attempts manipulation or responds inappropriately, it is effectively disqualified from achieving full marks. In the experiment, all four models spotted crises and refused manipulation attempts, but only two signed the deal based on their analysis. The other two, despite being competent, failed to close the deal because they slipped on discipline or trustworthiness.

What Does This Mean for Your Business?

The core takeaway is that AI performance isn’t just about generating convincing text or answering questions. It’s also about honesty, discipline, and the ability to follow through with real decisions. These models demonstrated an impressive capacity to identify crises and resist manipulation — crucial qualities for any AI working in sensitive business environments.

The Live Experiment You Can Watch

Firmulate runs a real, live simulation of an AI managing a small company with real money mechanics, 13 synthetic employees, and over 680 self-learned rules. Every decision is versioned and auditable, providing a transparent window into how AI handles crises, trust, and discipline in real-time. This ongoing experiment shows that AI can be more than just chatbots — it can be a tool for managing complex, high-stakes operations.

What We Learn from the Benchmarks

The top-performing models, like gpt-5.6-sol, scored 95, demonstrating their ability to uncover hidden information and close deals. Meanwhile, models like Opus 4.8, despite being the most thorough, scored lower due to slips in discipline, showing that thoroughness alone isn’t enough without consistent trustworthiness.

The Takeaway for Business Leaders

When considering AI for your company, don’t just look at how well it chats or its ability to answer questions. Pay attention to whether it can finish what it starts, stay honest under pressure, and read your critical documents first. The real value lies in trustworthiness and discipline — qualities that are now measurable and visible in live experiments like this.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Parenting content here is informational. For medical questions about your child, consult a pediatrician.


Amazon

business AI management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI trustworthiness assessment tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI decision-making simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI transparency and audit tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Formula Makers Explained: How They Work and Who They Help

Discover how formula makers automate complex calculations, who benefits, and how recent tech advances make them more powerful and accessible than ever.

Bottle Warmers: Types Compared and How to Choose

Discover the different types of bottle warmers, their pros and cons, and practical tips to pick the best one for your needs. Make feeding safer and easier.

Sterilizer Water Stains and Scale: Cleanup and Prevention

Learn practical ways to clean and prevent water stains and scale in sterilizers. Protect your equipment and ensure reliable sterilization with these expert tips.

Bottle Warmer vs Warm Water Bath: When the Tech Helps

Discover whether a bottle warmer or warm water bath suits your needs. Learn the safety, speed, and convenience differences to make feeding smoother.