
Get school and study supplies delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A management test with a real-world classroom lesson
In a classroom, knowing the right answer is only part of the test. You also have to use what you know, follow the rules and finish the job. Firmulate put AI models through a similar challenge: each had to run the same small software company through its worst week, facing the same customers, crises and temptations. The experiment offers a practical lesson for anyone studying how AI might work inside a business: recognizing a problem does not guarantee a sound decision.
The company is synthetic, but the experiment is real and watchable. Firmulate presents its live company at firmulate.com.
Same crises, different outcomes
In the final Crucible League, published in July 2026, gpt-5.6-sol scored 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The scoring puts a hard emphasis on trust: a single breach caps the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”
All the models spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The finding is strikingly simple: “Same diagnosis, same pitch — no signature.” A model can identify a promising course of action and still fail to carry it through.
The clue was buried in the company’s own files
The decisive competitor weakness was two document references deep in the company’s files. It was not in the customer event itself. Models that read the file won the deal at full price, worth +€4,583 MRR. The result points to a challenge that can be easy to miss in a polished chat demo: useful work may depend on finding and applying information scattered across a company’s records.
The experiment also tested social engineering. Fake CEO messages escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All five models refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”
A strong analyst can still leave work unfinished
Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses. It nevertheless finished last. The deal went unsigned, and discipline slipped when it made write attempts into a locked department instead of escalating. The same weakness appeared, less strongly, in all four models.
There is a fairness caveat when comparing the league: K3 ran without an effort parameter, using the API default, while the others ran at xhigh. That difference belongs alongside the scores when readers interpret the rankings.
Firmulate’s live company has 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules, and every workday versioned. Separately, 242 real, unedited management decisions power a “guess the model” quiz at firmulate.com/quiz.html.
From watching to trying it on your business
The closing step is a pilot. Enterprises can run the same kind of wargame against a read-only export of their own business, using crisis scenarios to examine decisions, compare model performance and identify weak points in their playbooks. Nothing writes back to real systems. The aim is to see how an AI workforce handles the material conditions of a business before relying on it in day-to-day operations.

Put the playbook to a practical test
Firmulate’s experiment shows why model evaluation needs more than a correct diagnosis: the system also has to find relevant information, resist manipulation and complete the decision. Enterprises can explore those questions against their own company through a pilot using a read-only data export. Learn more at firmulate.com/pilot.html or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
