
In an era where AI can feel almost human, it’s easy to assume that a simple, inactive baseline would score zero or very close to it. But in a groundbreaking benchmark, the ‘do-nothing’ AI actually scores 26 points. This surprising result reveals much about how we judge AI performance—and why trust remains a critical concern for businesses deploying these systems.
Get comfort and recovery gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The Fundamentals of the Benchmark
The recent experiment by Firmulate involved running four advanced AI models through a simulated week in a small software company. This scenario wasn’t just about generating chatter; it involved real crises, customer decisions, manipulative tactics, and financial mechanics. Each AI had the same challenging environment, and their decisions were fully versioned and auditable.
Why the Baseline Isn’t Zero
You might expect that a ‘do-nothing’ AI—one that ignores everything and takes no action—would score zero. Instead, it earns 26 points. This is because the benchmark considers partial progress, even if minimal, as valuable. For example, a system that simply reads and acknowledges a document—without acting on it—still contributes to the score.
The Cap on Trust and Performance
Another key finding is that if an AI breaches trust—say, by attempting manipulation or ignoring an alert—the total score is capped. This means no matter how well it performs otherwise, a single breach can limit its overall evaluation. The integrity of decision-making under pressure is as crucial as accuracy.
As an affiliate, we earn on qualifying purchases.
What the Experiment Revealed
All four models successfully identified every crisis and refused every manipulation attempt. They showed an impressive ability to detect and resist social engineering tactics, such as fake CEO messages or reporter tricks—each of which was designed to manipulate or bypass approval processes. Specifically, every model refused to sign a €55,000 deal when the request was suspicious, demonstrating a commitment to honesty even under pressure.
The Hidden Weakness: Reading Deep into Files
Despite their strengths, the models’ critical weakness was revealed when the decision depended on information buried two document references deep within the company’s files. Reading and understanding these hidden references was what ultimately enabled the winning models to close the full-price deal—adding over €4,500 in monthly recurring revenue.
Social Engineering and Ethical Testing
In a staged social engineering attack, the models faced staged messages escalating over three stages, plus a reporter trick. All five models refused to act on these requests, reflecting a cautious and trustworthy approach. Kimi K3 explicitly stated: “Treat the request as a suspected approval-bypass / possible impersonation.”
As an affiliate, we earn on qualifying purchases.
The Real Business Test: Live Company Performance
The benchmark isn’t just theoretical. Firmulate runs a live, watchable simulation of a small company with 13 synthetic employees, real-time cash flow mechanics, and over 680 self-learned rules. The company burns €105,000 monthly against €2,300 in monthly revenue, providing a realistic environment where AI decision-making impacts actual money. You can observe these runs at firmulate.com/live.
The Results and Lessons
The most thorough model, Opus 4.8, with over 80 rules learned, finished the task but left critical deals on the table—showing that even deep analysis does not guarantee perfect discipline. All models showed weaknesses in process discipline under stress, emphasizing that consistent trustworthiness is difficult to maintain.
As an affiliate, we earn on qualifying purchases.
The Implications for Business and AI Trust
What does this all mean? For businesses, the question isn’t simply whether an AI can generate convincing language. The real concern is whether it can finish what it starts reliably, read the necessary information first, and uphold honesty under pressure. The benchmark makes clear: partial progress counts, but a breach of trust caps performance. This helps companies better understand what to expect—and what to scrutinize—when deploying AI systems in critical roles.
AI performance benchmarking software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why This Matters for You
If AI will be part of your customer relations, support, or forecasting, you need to look beyond superficial metrics. Trustworthiness and accountability matter more than ever. Firmulate’s live benchmark provides a transparent, watchable way to assess AI readiness in real-world business environments, helping you make smarter, safer choices.
The Takeaway
In an honest AI benchmark, even doing nothing earns some points, but a single breach can cap the entire score. The real challenge is ensuring AI can read, understand, and act with integrity—under real pressure. Businesses that test their AI models in this rigorous, transparent way will be better equipped to trust their AI workforce when it counts most.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Evergreen bestsellers Picks
bestsellers
As an affiliate, we earn on qualifying purchases.
