firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Parents know the difference between a confident promise and a dependable decision. An AI assistant might sound calm while planning a family holiday, managing household bills or helping a small business serve its customers. But what happens when plans unravel, someone applies pressure and the next choice has real consequences?

For listenersOffer from Amazon

Turn the school run and nap time into listening time

  • Thousands of audiobooks, podcasts and originals
  • Listen on your phone, tablet or Echo — also offline
  • Cancel anytime
Try Audible free Free trial for new members
As an affiliate, we earn on qualifying purchases.

That question is at the heart of Firmulate, a live experiment that puts AI models in charge of a small software company. Its lesson for families and businesses alike: watching an AI talk is not the same as seeing how it handles a difficult week.

A company under pressure

In the final Crucible League, published in July 2026, frontier models faced the same small company, customers, crises and temptations. Each decision was versioned and auditable. The leaderboard placed gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The league’s trust principle is blunt: “no amount of good work outweighs a breach of trust.”

The surprising result was not that the models missed the danger. Every model spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. They reached the same diagnosis and delivered the same pitch; most still left the signature on the table.

The detail hiding in the files

The deal hinged on a competitor weakness buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. It is a practical reminder that a system can identify the broad problem and still fail to act on the evidence that matters.

Firmulate also tested pressure designed to bypass normal judgment: fake messages from a CEO escalating over three stages, followed by a reporter asking “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

One participant complicates the idea that more analysis automatically means better management. Opus 4.8 was the most thorough, with +80 learned rules and the deepest analyses, yet finished last. It left the close on the table and tried writing into a locked department instead of escalating. The same weakness appeared, more weakly, in all four models.

From watching to testing your own business

The public experiment is built around a live company with 13 synthetic employees and real money mechanics: burn of €105k/month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules and every workday versioned. Readers can watch the company at firmulate.com. A quiz built from 242 real, unedited management decisions invites visitors to guess which model made each call.

For enterprises, the next step is a pilot against their own business. Firmulate uses a read-only export to create a digital twin, then runs crisis scenarios to produce a board report with model rankings and weak points in the company’s playbooks. The boundary matters: nothing writes back to real systems.

There is a fairness caveat in the league. Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. The rankings are a record of this experiment, not a guarantee about every model or every workplace.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put your own playbooks to the test

The experiment shows why an AI evaluation needs more than polished answers: it should reveal whether a system can carry its own analysis through to action, find relevant evidence and hold its ground under pressure. A pilot lets a company examine those questions using its own business data while keeping the live systems untouched. Explore a Firmulate pilot or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Parenting content here is informational. For medical questions about your child, consult a pediatrician.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Feeding Tracker Apps: Logging Without Losing Your Mind

Discover how to make feeding tracking easier with simple tips, smart features, and realistic expectations. Keep your sanity while monitoring your child’s feeding habits.

When to Stop Sterilizing Bottles: A Practical Timeline

Learn the right age and circumstances to stop sterilizing bottles. Get a clear timeline backed by health guidance for safe, confident feeding practices.

Formula Maker Cleaning: A Realistic Weekly Routine

Discover practical tips for cleaning your formula maker weekly. Keep baby safe with a straightforward routine that balances hygiene and convenience.

AI’s Hidden Strength: Why Some Models Close Deals Under Pressure While Others Don’t

A live AI experiment reveals that only some models can close deals and stay honest under pressure, proving that true AI strength is measured beyond chat demos.