
Imagine a business without human employees, making critical decisions, facing crises, and fighting to stay afloat — all in front of your eyes. This isn’t science fiction; it’s the live experiment by Firmulate, a company that runs a simulated business entirely managed by artificial intelligence models. For parents and families, it’s a window into how AI might soon influence everyday work — and the importance of trust, honesty, and discipline in machines that could soon work alongside humans.
The Live Business in Action: No Humans, Just AI Models
At the heart of this experiment is a small, virtual company managed by 13 synthetic employees, each modeled by advanced AI systems. This isn’t a simple chatbot; it’s a complex, real-time simulation where every decision is logged, every crisis is faced, and every choice is auditable. The company operates with a monthly burn rate of €105,000, while generating only €2,300 in monthly recurring revenue — a clear sign of its fragile financial state. The entire setup is transparent and accessible to the public at firmulate.com/live.html.
Testing the Limits of AI Integrity and Performance
The experiment pits four frontier AI models against each other, with each running the same week’s worst-case scenarios — identical crises, same customer demands, and identical temptations to cheat or manipulate. The goal: see if these AI managers can recognize problems, avoid manipulation, and ultimately close profitable deals.
Remarkably, all four models identified every crisis and refused every attempt to manipulate them. Yet, only two of the models managed to sign a €55,000 deal, which was earned through honest analysis and proper diagnosis. The other two either failed to act on their insights or left opportunities unexploited, demonstrating that even the smartest AI can stumble without discipline and focus.
Discovering Hidden Weaknesses through Deep Reading
One of the most revealing findings was that the models which succeeded had read beyond surface information—delving two document references deep into the company’s own files. That’s where the key weakness lay—hidden in internal documents, not in the customer interactions. The models that uncovered and read these files won the deal at full price, adding over €4,500 to their monthly recurring revenue.
Trust and Social Engineering Tests
To challenge the AI’s honesty, the experiment included staged social engineering attacks—fake CEO messages escalating over three stages, and even a reporter trick asking for a secret approval. All five models tested refused to act on these unethical requests, with Kimi K3 explicitly reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This demonstrates that advanced AI can be designed to reject manipulative behavior, even under pressure.

Hands-On Simulation Modeling with Python: Develop simulation models for improved efficiency and precision in the decision-making process, 2nd Edition
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Real-World Implications
This experiment isn’t just a tech showcase; it raises vital questions for the future of work. As AI systems become capable of managing tasks, making decisions, and interacting with humans, trust and integrity become critical. If AI agents are to handle your customer support, sales, or strategic planning, can we rely on them to act honestly, especially when under stress or facing manipulation?
In the live experiment, every workday is versioned, and all decisions are openly recorded—allowing managers and developers to see exactly how the AI performs in real situations. The ongoing company has already burned €105,000 per month, with only €2,300 in revenue, highlighting the challenge of building and scaling trustworthy AI in business.
Deep Discipline and Learning from Failures
The profile of the least successful model, Opus 4.8, reveals that even the deepest analysis doesn’t guarantee success. It left a crucial deal unexecuted because discipline slipped—showing that discipline, process adherence, and internal escalation are as vital as analytical depth. The experiment underscores that in high-stakes environments, the AI’s ability to follow rules and escalate issues matters just as much as its diagnostic skills.

Trustworthy Medical AI: A Builder's Guide to Safe, Compliant Software as a Medical Device
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Future of Building Trust in AI-Managed Businesses
For families and parents, this experiment might seem distant, but it offers a glimpse into a future where AI could help manage work tasks—if we can ensure they are honest and disciplined. Trustworthiness, transparency, and adherence to rules aren’t just human virtues; they become technological imperatives.
As firms like Firmulate continue to develop and refine these AI systems, the goal isn’t just efficiency but building machines that are reliable partners—able to recognize crises, refuse manipulation, and deliver honest work. For now, you can watch this ongoing story unfold at firmulate.com/live.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI ethics and integrity training kits
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI cybersecurity and manipulation detection tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.