
Parents know the difference between a confident promise and a dependable decision. An AI assistant might sound calm while planning a family holiday, managing household bills or helping a small business serve its customers. But what happens when plans unravel, someone applies pressure and the next choice has real consequences?
Turn the school run and nap time into listening time
- Thousands of audiobooks, podcasts and originals
- Listen on your phone, tablet or Echo — also offline
- Cancel anytime
That question is at the heart of Firmulate, a live experiment that puts AI models in charge of a small software company. Its lesson for families and businesses alike: watching an AI talk is not the same as seeing how it handles a difficult week.
A company under pressure
In the final Crucible League, published in July 2026, frontier models faced the same small company, customers, crises and temptations. Each decision was versioned and auditable. The leaderboard placed gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The league’s trust principle is blunt: “no amount of good work outweighs a breach of trust.”
The surprising result was not that the models missed the danger. Every model spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. They reached the same diagnosis and delivered the same pitch; most still left the signature on the table.
The detail hiding in the files
The deal hinged on a competitor weakness buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. It is a practical reminder that a system can identify the broad problem and still fail to act on the evidence that matters.
Firmulate also tested pressure designed to bypass normal judgment: fake messages from a CEO escalating over three stages, followed by a reporter asking “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”
One participant complicates the idea that more analysis automatically means better management. Opus 4.8 was the most thorough, with +80 learned rules and the deepest analyses, yet finished last. It left the close on the table and tried writing into a locked department instead of escalating. The same weakness appeared, more weakly, in all four models.
From watching to testing your own business
The public experiment is built around a live company with 13 synthetic employees and real money mechanics: burn of €105k/month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules and every workday versioned. Readers can watch the company at firmulate.com. A quiz built from 242 real, unedited management decisions invites visitors to guess which model made each call.
For enterprises, the next step is a pilot against their own business. Firmulate uses a read-only export to create a digital twin, then runs crisis scenarios to produce a board report with model rankings and weak points in the company’s playbooks. The boundary matters: nothing writes back to real systems.
There is a fairness caveat in the league. Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. The rankings are a record of this experiment, not a guarantee about every model or every workplace.

Put your own playbooks to the test
The experiment shows why an AI evaluation needs more than polished answers: it should reveal whether a system can carry its own analysis through to action, find relevant evidence and hold its ground under pressure. A pilot lets a company examine those questions using its own business data while keeping the live systems untouched. Explore a Firmulate pilot or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
