
Imagine testing an employee who does nothing but still manages to score 26 out of 100 on their performance review. That’s the reality of the latest AI benchmark, revealing what’s behind the glowing scores and what truly matters in AI management. For those in the world of quality and trust—think bakers ensuring every ingredient counts—this experiment underscores critical lessons on reliability and integrity in AI systems.
Get baking supplies delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The Honest Face of AI Benchmarks
In the rapidly evolving AI landscape, benchmarks are central to understanding how models perform in real-world scenarios. But what if a model, doing nothing, still scores 26? That’s not a mistake—it’s a deliberate baseline set by recent testing conducted by Firmulate, an innovative platform that simulates the management of a small software company through its worst week. This baseline score of 26 points is revealing because it captures the essence of a ‘do-nothing’ approach, where partial progress is acknowledged, but trustworthiness remains paramount.
Why Partial Progress Counts
In this experiment, every decision made by the AI models was documented and auditable. Remarkably, all models identified every crisis and refused every manipulation attempt—an essential measure of honesty. Yet, only two of the four models managed to successfully close a deal worth €55,000, even when their analysis identified the opportunity. This illustrates a key insight: in trustworthy AI, doing the right thing consistently is more valuable than just identifying opportunities or passing tests.
The Significance of a Single Breach
The scoring system caps at 26 points for the ‘do-nothing’ baseline, and it’s designed to reflect trust. A single breach—such as signing a deal without proper validation—caps the score, emphasizing that even the smallest lapse can undermine an AI’s reliability. This rule underscores a simple truth: no amount of good performance can outweigh a breach of trust. In real-world business, integrity is non-negotiable, just like a baker’s commitment to quality ingredients.
AI trustworthiness testing tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What the Experiment Tells Us About AI in Business
The experiment involved running four frontier AI models against the same challenging scenario, simulating a small company’s tough week: customer crises, internal crises, manipulation attempts, and ethical dilemmas. Each model faced identical conditions, and decision-making was meticulously versioned and auditable, ensuring transparency.
Trust Under Pressure
All models demonstrated competence by recognizing every crisis and refusing manipulative tactics, such as fake CEO messages and reporter tricks. For instance, when faced with staged social engineering attempts—escalating messages and background requests—every model refused to comply, citing concerns over impersonation and bypassing approval processes. This shows that, even amid pressure, AI models can prioritize honesty.
Reading Beyond the Surface
One of the critical weaknesses uncovered was in models’ ability to process deeper information. The decisive factor was whether the AI read and understood documents stored two levels deep within the company’s files. The models that successfully read these internal files directly won the deal at full price (€4,583 MRR), illustrating that thoroughness and attention to detail can be decisive in business negotiations.
Performance and Discipline
Among the models tested, Opus 4.8 stood out for its thoroughness—an analysis with over 80 learned rules. Yet, it finished last because it left the close on the table and showed signs of slipping discipline, such as attempting to escalate issues into a locked department rather than following proper procedures. This exemplifies that in AI and in business, depth must be matched with disciplined execution.
AI transparency and audit software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why the Benchmarks Matter for Business Leaders
For managers, entrepreneurs, and decision-makers, this experiment offers a vital lesson: AI’s true value lies not merely in its ability to generate text or simulate conversation but in its capacity to stay honest, read deeply, and follow through reliably. The scores—ranging from 95 for the top-performing GPT-5.6 to 77 for another model—highlight that even the best models occasionally slip, but trustworthiness remains the core metric.
Running Your Own Wargame
Businesses can test their AI tools through a similar simulated environment, known as a wargame, which runs a copy of their operations without risking real systems. This allows companies to evaluate how their AI agents handle crises, manipulative tactics, and internal procedures—crucial for ensuring trustworthy AI deployment in sensitive areas like customer support, sales, and operations.
As an affiliate, we earn on qualifying purchases.
In Conclusion: The Bottom Line for AI Trustworthiness
In the end, the experiment confirms a profound truth: partial progress and high scores are not enough if the AI system breaches trust. A do-nothing baseline scoring 26 points reminds us that integrity, thoroughness, and discipline are foundational—especially when AI touches business-critical functions. For those baking the future of AI into their operations, it’s a call to prioritize trustworthy performance over flashy metrics, ensuring that every decision, like every ingredient, is carefully considered.

Trustworthiness in AI isn’t just about high scores or quick wins. It’s about integrity, thoroughness, and consistent discipline—basics that even a do-nothing baseline score highlights as essential for reliable performance in business.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
