
Imagine buying a new designer dress and discovering it barely fits or falls apart the moment you wear it. You’d feel cheated, right? The same goes for AI benchmarks—what looks like progress can sometimes mask fundamental flaws. In a recent experiment, even a do-nothing AI baseline scored a surprising 26 out of 100, revealing crucial insights about how we measure AI’s true readiness for prime time.
Get your wardrobe favorites delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
When evaluating AI systems for business-critical tasks, the stakes are high. A recent public experiment conducted by Firmulate puts this into sharp focus by testing multiple AI models as if they were running a small software company facing its worst week. Every decision was documented and auditable, and the results challenge common assumptions about what these models can actually do.
The experiment involved four leading AI models navigating the same simulated environment: a company with real customers, crises, and temptations to cheat. The models had clear instructions: spot every crisis, refuse manipulation, and close deals when appropriate. The results? All four models identified every crisis and refused every manipulation attempt, showing they understand the rules of engagement. Yet only two successfully signed off on a €55,000 deal, earning full points for their analysis.
Why did the other two fall short? The key lies in a buried fact within the company’s own files—information two document references deep—that the winning models read and used to close the deal. The losing models failed to dig this deep, illustrating that reading and understanding context is vital for real-world success. This subtle flaw underscores a critical point: surface-level performance, like chat responses, can be deceiving. True competence requires depth and diligence—traits that are often overlooked in standard benchmarks.
Another telling detail involves social engineering attempts, such as fake CEO messages and reporter tricks. All models refused these phishing tactics, with Kimi K3 explaining: “Treat the request as a suspected approval-bypass / possible impersonation.” This indicates a baseline level of trustworthiness that all models maintained, an essential trait when AI interacts with humans in sensitive roles.
However, the results also revealed a stark reality: even the most thorough model, Opus 4.8, failed to close the deal fully. Despite analyzing over 80 rules and providing deep insights, it left the close on the table and slipped into process slips—like writing attempts into a locked department instead of escalating them. This highlights that even deep analysis doesn’t guarantee flawless execution, especially under real-world pressures.
The entire experiment underscores a vital point for businesses: scores alone don’t tell the full story. A benchmark that assigns a minimum score of 26 to a do-nothing baseline reveals the importance of honesty and transparency in evaluation. It’s a reminder that partial progress counts and that trust breaches cap the maximum score of any AI system, reflecting real-world limitations.
For companies considering AI as a strategic partner—whether in customer support, decision-making, or operational management—the key takeaway is clear: performance must go beyond surface-level chat and include depth, diligence, and integrity. Firms should run their own ‘wargames’ against AI systems before deploying them in critical roles, ensuring they can handle crises, read complex documents, and stay honest under pressure.
Visit Firmulate’s live platform to see real-time experiments and understand how AI models perform in simulated business environments. These insights help clarify what a true, trustworthy AI workforce looks like—one that can handle real crises, read your files, and close deals responsibly.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
business AI decision support systems
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
