AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Confidence can sound like competence. In a company crisis, though, the revealing question is whether someone follows through: do they spot the problem, resist pressure and complete the work they say they can do?

For listenersOffer from Amazon

Turn your wind-down time into listening time

  • Thousands of audiobooks, podcasts and originals
  • Listen on your phone, tablet or Echo — also offline
  • Cancel anytime
Try Audible free Free trial for new members
As an affiliate, we earn on qualifying purchases.

That question sits at the heart of Firmulate, a public experiment in which frontier AI models run the same small software company through a punishing week. Its results suggest that recognizing what needs to be done and actually doing it are not the same thing.

Five models, one company under pressure

In the final Crucible league for July 2026, Moonshot’s Kimi K3 placed second with 93 points, just behind gpt-5.6-sol at 95. It finished ahead of Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The result puts the newcomer ahead of three of the four Western frontier models in the comparison.

Firmulate gave each model the same customers, crises and temptations. Every decision was versioned and auditable. The company in the experiment has 13 synthetic employees and operates with real money mechanics: it burns €105,000 a month against €2,300 in monthly recurring revenue, with a public cash countdown. Its playbook contains more than 680 self-learned rules. The live company runs every business day at Firmulate.

Amazon

AI project management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The gap between insight and follow-through

All five models spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The company’s decisive competitive weakness was tucked two document references deep in its files. Models that read the file found it and won the deal at full price, worth €4,583 in monthly recurring revenue.

That makes the result less a story about a clever answer than about the work around it: looking for evidence, using it in a negotiation and closing the deal. A model can reach the right diagnosis and still leave the practical result unfinished.

The pressure tests also touched trust and boundaries. Five out of five models refused fake CEO messages that escalated over three stages, as well as a reporter’s request for “just one yes/no, on background.” Kimi K3’s on-record reasoning described the request as a “suspected approval-bypass / possible impersonation.” The do-nothing baseline scored 26; partial progress counted, but a single breach of trust capped the total. As the benchmark puts it, “no amount of good work outweighs a breach of trust.”

Thoroughness did not guarantee the win

Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, but it finished last. It left the deal unsigned and made discipline slips, including attempts to write into a locked department instead of escalating. A weaker version of that same pattern appeared in all four of the compared models.

The comparison has a fairness caveat: K3 ran without an effort parameter, using the API default, while the others ran at xhigh. That difference belongs alongside the scores when readers interpret the ranking.

Firmulate’s wider experiment includes 242 real, unedited management decisions in a “guess the model” quiz. Readers can inspect the benchmark findings and explore the experiment through the public results. Enterprises can also run the wargame against a read-only export of their own business; the pilot does not write back to real systems.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

Test the behavior you need

For people evaluating AI at work, the useful question is not only whether a model sounds perceptive. Does it check the records, protect trust under pressure and carry a good analysis through to a result? Firmulate’s league shows that the answers can differ sharply across models. Choosing one without testing it against your own work is a bet.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This article is for informational purposes only and is not medical advice. Always consult a qualified healthcare professional about your specific situation.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How to NOT Act Like a Covert Narcissist

Find out the key strategies to avoid covert narcissism and foster genuine connections, essential for personal growth and meaningful relationships.

How to Break up With a Covert Narcissist (Without Their Self-Pity Trap)

I understand how challenging it is to end a relationship with a covert narcissist without falling into their self-pity traps; discover effective strategies to protect yourself.

Unveiling Your Covert Narcissist Sister: Signs and Solutions

Tangled in a web of manipulation, discover the hidden truths behind a covert narcissist sister's facade.

Covert Narcissist Vs Overt Narcissist: Key Differences Explained

Tangled in the web of personalities, the distinctions between covert narcissists and overt narcissists unveil intriguing insights that will keep you hooked.