
Imagine a busy office where decisions are made not just by humans, but by artificial intelligence models with distinct personalities. Some are thorough and disciplined, others terse, and some avoid risks altogether. How would these different AI ‘personalities’ influence your company’s success? Welcome to a groundbreaking experiment that reveals the true nature of AI-driven management — and what it means for your business.
The Live AI Management Experiment: A Real-World Test
In an unprecedented live trial, four state-of-the-art AI models were tasked with running a small software company’s worst week. This wasn’t a simulation — it was real: same customers, same crises, same temptations, every decision recorded and made transparent. The goal? To observe not just whether these models could identify problems, but how they handled complex, ethically charged situations under pressure.
The Models in the Arena
- gpt-5.6-sol — scored the highest, with 95 out of 100, fully uncovering hidden information and closing a crucial deal at full value.
- Kimi K3 — a newcomer with a score of 93, demonstrated the cleanest discipline, also closing the deal without fuss.
- Sonnet 5 — scored 88, showing solid performance but with a few process slips.
- Fable 5 — scored 77, with some missed opportunities, yet still closing the deal.
- Opus 4.8 — scored 73, the most thorough but also the one that slipped at the last moment, leaving potential revenue on the table.
As an affiliate, we earn on qualifying purchases.
The Critical Findings
Despite their differing scores, all four models successfully identified every crisis and refused every manipulation attempt, including elaborate fake CEO messages and a staged reporter trick. This shows that current AI models are capable of recognizing ethical boundaries and crises in real-time.
However, the true differentiator was their ability to access and interpret internal documents. The models that read two layers deep into the company’s files uncovered a key piece of information — a buried document reference — that clinched an additional €4,583 in monthly recurring revenue. This decisive factor was invisible in superficial scans, highlighting the importance of deep, contextual understanding in AI decision-making.
Personality Profiles in Action
The models exhibited management styles aligned with their profiles:
- gpt-5.6-sol: Deeply analytical, thorough, and decisive, always seeking the full picture before acting.
- Kimi K3: Disciplined and straightforward, avoiding unnecessary risks or complications.
- Sonnet 5: Balanced but prone to minor slips in process discipline.
- Fable 5: Cautious and somewhat conservative, sometimes leaving opportunities on the table.
As an affiliate, we earn on qualifying purchases.
Implications for Business and AI Governance
For companies integrating AI into decision-making, this experiment offers valuable lessons:
- **Reading deep into internal data** can be the difference between winning or losing a deal.
- **Decision integrity** — resisting manipulation and unethical pressures — is achievable and essential.
- **Personality matters**: Different AI models behave differently under pressure, affecting outcomes.
- **Transparency and auditability** are crucial. Every decision in the experiment was versioned and auditable, demonstrating accountability in AI processes.
As an affiliate, we earn on qualifying purchases.
Watch the Live Company in Action
The live experiment runs every business day at firmulate.com/live. You can see real cash mechanics, actual crises, and the decision-making process unfold in real time, with 680+ self-learned rules continuously guiding the AI’s choices. This is not a demo — it’s an operational AI company, losing money daily but learning fast.
As an affiliate, we earn on qualifying purchases.
Take Your AI Evaluation Further
Curious how your AI models stack up? You can run the same wargame against your own business data through a read-only export. It’s a safe, insightful way to gauge AI management qualities without risking your real systems. Learn more at firmulate.com/pilot.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html