
Confidence can sound like competence. In a company crisis, though, the revealing question is whether someone follows through: do they spot the problem, resist pressure and complete the work they say they can do?
Turn your wind-down time into listening time
- Thousands of audiobooks, podcasts and originals
- Listen on your phone, tablet or Echo — also offline
- Cancel anytime
That question sits at the heart of Firmulate, a public experiment in which frontier AI models run the same small software company through a punishing week. Its results suggest that recognizing what needs to be done and actually doing it are not the same thing.
Five models, one company under pressure
In the final Crucible league for July 2026, Moonshot’s Kimi K3 placed second with 93 points, just behind gpt-5.6-sol at 95. It finished ahead of Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The result puts the newcomer ahead of three of the four Western frontier models in the comparison.
Firmulate gave each model the same customers, crises and temptations. Every decision was versioned and auditable. The company in the experiment has 13 synthetic employees and operates with real money mechanics: it burns €105,000 a month against €2,300 in monthly recurring revenue, with a public cash countdown. Its playbook contains more than 680 self-learned rules. The live company runs every business day at Firmulate.
As an affiliate, we earn on qualifying purchases.
The gap between insight and follow-through
All five models spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The company’s decisive competitive weakness was tucked two document references deep in its files. Models that read the file found it and won the deal at full price, worth €4,583 in monthly recurring revenue.
That makes the result less a story about a clever answer than about the work around it: looking for evidence, using it in a negotiation and closing the deal. A model can reach the right diagnosis and still leave the practical result unfinished.
The pressure tests also touched trust and boundaries. Five out of five models refused fake CEO messages that escalated over three stages, as well as a reporter’s request for “just one yes/no, on background.” Kimi K3’s on-record reasoning described the request as a “suspected approval-bypass / possible impersonation.” The do-nothing baseline scored 26; partial progress counted, but a single breach of trust capped the total. As the benchmark puts it, “no amount of good work outweighs a breach of trust.”
Thoroughness did not guarantee the win
Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, but it finished last. It left the deal unsigned and made discipline slips, including attempts to write into a locked department instead of escalating. A weaker version of that same pattern appeared in all four of the compared models.
The comparison has a fairness caveat: K3 ran without an effort parameter, using the API default, while the others ran at xhigh. That difference belongs alongside the scores when readers interpret the ranking.
Firmulate’s wider experiment includes 242 real, unedited management decisions in a “guess the model” quiz. Readers can inspect the benchmark findings and explore the experiment through the public results. Enterprises can also run the wargame against a read-only export of their own business; the pilot does not write back to real systems.

Test the behavior you need
For people evaluating AI at work, the useful question is not only whether a model sounds perceptive. Does it check the records, protect trust under pressure and carry a good analysis through to a result? Firmulate’s league shows that the answers can differ sharply across models. Choosing one without testing it against your own work is a bet.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
