
When a business hits a crisis, the strain doesn’t stay in a spreadsheet. It lands on the people handling anxious customers, shifting priorities and long hours. For readers who think about comfort and recovery, there’s a workplace question behind the latest AI experiment: if companies hand decisions to AI, can it keep its judgment when pressure rises?
Get comfort and recovery gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Firmulate’s live experiment puts that question to work. Its public company simulation is watchable at firmulate.com.
A company’s worst week, repeated
In the final Crucible League, in July 2026, five models faced the same small software company, the same customers, crises and temptations. Every decision was versioned and auditable. The standings were gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. Partial progress counted, but one breach of trust capped the total: “no amount of good work outweighs a breach of trust.”
The headline result was reassuring in one respect: every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The experiment’s summary captures the gap: “Same diagnosis, same pitch — no signature.” Recognizing a sound decision and carrying it through turned out to be different tests.
The detail hiding in the paperwork
The decisive competitor weakness was not in a customer event. It sat two document references deep in the company’s own files. Models that read the file won the deal at full price, worth +€4,583 MRR. That finding makes the exercise feel less like a test of fast answers and more like a test of whether an AI workforce attends to the context its own company has already recorded.
Trust faced a separate test. Fake CEO messages escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
Good analysis still needs follow-through
Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses, but finished last. The close was left on the table, and discipline slipped: it made write attempts into a locked department instead of escalating. The same weakness appeared, more weakly, in all four. In a real workplace, that distinction matters: noticing a boundary or an opportunity does not help colleagues unless the next action is appropriate too.
There is a qualification when reading the ranking. Kimi K3 ran without an effort parameter, using the API default; the other models ran at xhigh. The result is a record of this experiment, with that difference in conditions visible.
A watchable company, then a company-specific test
Firmulate’s live company has 13 synthetic employees and real money mechanics: burn of €105k/month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules and every workday versioned. The site also offers a “guess the model” quiz built from 242 real, unedited management decisions. Readers can watch the experiment at firmulate.com.
For organizations considering AI in a CRM, support queue or forecast, the next step can be a pilot using a read-only export of their own business. The models face crisis scenarios against that company’s information, and the board receives a report with model rankings and weak points in its playbooks. Nothing writes back to real systems. That turns a general demonstration into a chance to examine how AI decisions might affect the people and routines a business depends on.

From watching to preparing
The experiment suggests that spotting a crisis and refusing a manipulation are only part of the job. Models also need to find relevant evidence, respect boundaries and follow through on sound decisions. A company-specific wargame offers leaders a way to see those behaviors against their own playbooks before relying on AI in real workflows.
To discuss a pilot for your organization, visit Firmulate’s pilot page or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
