
Fashion runs on sharp instincts: spotting what customers want, reading a shifting market and knowing when a promising opportunity is ready to close. But an AI assistant can identify the right move and still fail to make it. In Firmulate’s business experiment, every model spotted every crisis. Only two signed the €55,000 deal their own analysis had earned.
Get your wardrobe favorites delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
From polished answers to pressure-tested decisions
Chat can make an AI look decisive. Running a company through a difficult week is a different test. Firmulate gave each frontier model the same small software company, the same customers, the same crises and the same temptations. Every decision was versioned and auditable.
The final Crucible League, in July 2026, ranked gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The experiment’s standard was deliberately unforgiving: “no amount of good work outweighs a breach of trust.”
The quiet cost of hesitation
All the models found every crisis and refused every manipulation attempt. Yet only two completed the sale their own analysis had justified. The gap was not in recognizing the opportunity or making the pitch. It was in following through: “Same diagnosis, same pitch — no signature.”
The decisive competitor weakness was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. It is a pointed lesson for any business considering AI: a system can sound perceptive while overlooking the detail that changes the outcome.
Trust under pressure, discipline under strain
The test also included fake CEO messages escalating over three stages and a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 explained its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”
Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, yet it finished last. It left the close on the table and discipline slipped: it tried to write into a locked department instead of escalating. A weaker version of that same weakness appeared in all four models. Thoroughness, the experiment suggests, does not guarantee sound execution.
There is a fairness caveat in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. Readers can also judge the decisions themselves: 242 real, unedited management decisions power Firmulate’s “guess the model” quiz at firmulate.com.
A live company, then a company-specific test
Firmulate’s live company has 13 synthetic employees and real money mechanics: burn of €105k/month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules and every workday versioned. The live experiment is watchable at Firmulate. It makes the broader question concrete: how will an AI workforce behave when decisions touch customers, money and trust?
For enterprises, the next step is a pilot against a read-only export of their own business. The export can represent customers, pipeline and company rules; the wargame can put that business through crisis scenarios and produce a board report with model rankings and weak points in its playbooks. Nothing writes back to real systems.

Put your own playbooks to the test
Watching an AI company is a useful start. A pilot can show how models handle the pressures and opportunities specific to yours. Explore the Firmulate pilot and contact contact@firmulate.com to discuss a wargame for your business.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Evergreen bestsellers Picks
bestsellers
As an affiliate, we earn on qualifying purchases.
