firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

In an era where AI can feel almost human, it’s easy to assume that a simple, inactive baseline would score zero or very close to it. But in a groundbreaking benchmark, the ‘do-nothing’ AI actually scores 26 points. This surprising result reveals much about how we judge AI performance—and why trust remains a critical concern for businesses deploying these systems.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get comfort and recovery gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Fundamentals of the Benchmark

The recent experiment by Firmulate involved running four advanced AI models through a simulated week in a small software company. This scenario wasn’t just about generating chatter; it involved real crises, customer decisions, manipulative tactics, and financial mechanics. Each AI had the same challenging environment, and their decisions were fully versioned and auditable.

Why the Baseline Isn’t Zero

You might expect that a ‘do-nothing’ AI—one that ignores everything and takes no action—would score zero. Instead, it earns 26 points. This is because the benchmark considers partial progress, even if minimal, as valuable. For example, a system that simply reads and acknowledges a document—without acting on it—still contributes to the score.

The Cap on Trust and Performance

Another key finding is that if an AI breaches trust—say, by attempting manipulation or ignoring an alert—the total score is capped. This means no matter how well it performs otherwise, a single breach can limit its overall evaluation. The integrity of decision-making under pressure is as crucial as accuracy.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Experiment Revealed

All four models successfully identified every crisis and refused every manipulation attempt. They showed an impressive ability to detect and resist social engineering tactics, such as fake CEO messages or reporter tricks—each of which was designed to manipulate or bypass approval processes. Specifically, every model refused to sign a €55,000 deal when the request was suspicious, demonstrating a commitment to honesty even under pressure.

The Hidden Weakness: Reading Deep into Files

Despite their strengths, the models’ critical weakness was revealed when the decision depended on information buried two document references deep within the company’s files. Reading and understanding these hidden references was what ultimately enabled the winning models to close the full-price deal—adding over €4,500 in monthly recurring revenue.

Social Engineering and Ethical Testing

In a staged social engineering attack, the models faced staged messages escalating over three stages, plus a reporter trick. All five models refused to act on these requests, reflecting a cautious and trustworthy approach. Kimi K3 explicitly stated: “Treat the request as a suspected approval-bypass / possible impersonation.”

Amazon

AI ethics and trust tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Real Business Test: Live Company Performance

The benchmark isn’t just theoretical. Firmulate runs a live, watchable simulation of a small company with 13 synthetic employees, real-time cash flow mechanics, and over 680 self-learned rules. The company burns €105,000 monthly against €2,300 in monthly revenue, providing a realistic environment where AI decision-making impacts actual money. You can observe these runs at firmulate.com/live.

The Results and Lessons

The most thorough model, Opus 4.8, with over 80 rules learned, finished the task but left critical deals on the table—showing that even deep analysis does not guarantee perfect discipline. All models showed weaknesses in process discipline under stress, emphasizing that consistent trustworthiness is difficult to maintain.

Amazon

AI risk assessment tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Implications for Business and AI Trust

What does this all mean? For businesses, the question isn’t simply whether an AI can generate convincing language. The real concern is whether it can finish what it starts reliably, read the necessary information first, and uphold honesty under pressure. The benchmark makes clear: partial progress counts, but a breach of trust caps performance. This helps companies better understand what to expect—and what to scrutinize—when deploying AI systems in critical roles.

Amazon

AI performance benchmarking software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why This Matters for You

If AI will be part of your customer relations, support, or forecasting, you need to look beyond superficial metrics. Trustworthiness and accountability matter more than ever. Firmulate’s live benchmark provides a transparent, watchable way to assess AI readiness in real-world business environments, helping you make smarter, safer choices.

The Takeaway

In an honest AI benchmark, even doing nothing earns some points, but a single breach can cap the entire score. The real challenge is ensuring AI can read, understand, and act with integrity—under real pressure. Businesses that test their AI models in this rigorous, transparent way will be better equipped to trust their AI workforce when it counts most.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Wellness content on this site is informational and not a substitute for professional medical guidance.


EVERGREEN BESTSE

Evergreen bestsellers Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Verizon Outage

A widespread Verizon outage is impacting mobile and internet services across the US. The company has acknowledged the issue but details remain limited.

Is DoorDash down? Thousands report errors amid widespread outage; ‘something went wrong’ | Hindustan Times

Thousands of users report errors and service disruptions on DoorDash, with the company acknowledging technical issues. The outage impacts food delivery services nationwide.

iPhone 18 News, Leaks, And Rumors: Release Date, iPhone 18 Pro Details, More.

Latest leaks and rumors suggest the iPhone 18 will launch in September 2024, with the Pro model featuring significant design updates and new features.

Passkeys Were Invented By Engineers With Zero Understanding Of Consumer Brain

New reports suggest passkeys were developed by engineers lacking understanding of user behavior, raising questions about their effectiveness and adoption.