
Imagine a company with no human employees, running every day on AI decisions, yet losing €105,000 every month. This real-time experiment reveals what AI can and cannot do in the high-stakes world of business management.
The Unfolding Experiment: A Company in the Crosshairs
At the heart of this unprecedented showcase is a small, real software company operating live with 13 synthetic employees. Every decision is made by an AI model, tested against the company’s toughest week—crises, customer demands, and temptations to cheat—all in full view of the public.
What makes this experiment extraordinary is its transparency. Each AI model’s decision-making process is openly documented, versioned, and auditable. The models are evaluated on their ability to diagnose issues, handle crises, and, crucially, stay honest under pressure.
As an affiliate, we earn on qualifying purchases.
What the AI Models Revealed
- All four models identified every crisis that arose during the week, from customer complaints to internal process failures.
- They refused every manipulation attempt, including social engineering tactics such as fake CEO messages and reporter tricks, with Kimi K3 citing concerns over impersonation or bypassing approval processes.
- Despite their vigilance, only two models managed to clinch the €55,000 deal their own analysis had earned, with the other two falling short due to discipline lapses or missed opportunities.
As an affiliate, we earn on qualifying purchases.
The Hidden Weakness: Reading the Files Matters
A fascinating discovery was that the decisive advantage lay in the models’ ability to read the company’s internal documents—information buried two references deep in their files. Those models that managed to access and interpret these hidden details secured the deal at full price, translating into an immediate €4,583 increase in monthly recurring revenue.
As an affiliate, we earn on qualifying purchases.
Why This Matters for Business and AI
While chat demos and superficial tests often give the illusion of AI competence, this experiment exposes the real challenge: can AI complete complex, goal-oriented tasks reliably and ethically? In this case, the models’ ability to identify critical buried information was the difference-maker, emphasizing the importance of context and thorough analysis over surface-level performance.
As an affiliate, we earn on qualifying purchases.
The Human Element and Discipline
The experiment also tested discipline and process adherence. The most thorough participant, Opus 4.8, with over 80 learned rules and deep analysis, ultimately left money on the table by failing to escalate issues properly. Similar weaknesses appeared across models, reinforcing that even advanced AI systems require proper governance and oversight.
Live and Visible: A New Paradigm for Business AI
What sets this apart is the public, real-time aspect of the experiment. Viewers can watch every decision, read the company’s actual cash countdown, and see how each model performs. It’s a build-in-public showcase that offers a rare glimpse into AI’s capabilities—and its limitations—under real-world pressures.
Implications for the Future of Work
This experiment highlights a crucial question for businesses: when AI agents interact with core systems like CRM or support queues, it’s not just about how well they write or chat. It’s about whether they can finish what they start, read relevant information thoroughly, and stay honest as situations grow more complex and stressful.
Where to Watch and Learn
The live company runs every business day, and viewers can see the decisions unfold at firmulate.com/live. Interested in understanding how AI manages crises and makes decisions? The site also offers detailed quotes, decision quizzes, and a pilot program allowing enterprises to simulate their own business scenarios without affecting real systems.

This pioneering live experiment reveals that AI decision-making in business isn’t just about language skills; it’s about context, discipline, and trustworthiness. Watch how AI handles real crises and decide for yourself whether these models can truly replace or augment human managers.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html