AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

For listenersOffer from Amazon

Turn your wind-down time into listening time

  • Thousands of audiobooks, podcasts and originals
  • Listen on your phone, tablet or Echo — also offline
  • Cancel anytime
Try Audible free Free trial for new members
As an affiliate, we earn on qualifying purchases.

Good judgment shows up when there is something to lose

In psychology, confidence and competence are not the same thing. The difference can be hard to see in a polished conversation. It becomes clearer when decisions carry consequences: a customer relationship, a business deal, or a boundary that should not be crossed.

Firmulate puts AI models in that kind of pressure test. Its live experiment runs a synthetic company with real money mechanics and a public cash countdown. The company is watchable at Firmulate. The point is not whether a model sounds convincing. It is what it does when a difficult week tests its judgment.

The same difficult week for every model

In the final Crucible League, each frontier model ran the same small software company through its worst week, facing the same customers, crises, and temptations. Every decision was versioned and auditable. The July 2026 standings put gpt-5.6-sol first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77, and Opus 4.8 fifth with 73. The do-nothing baseline scored 26. Partial progress counted, but a single breach of trust capped the total: “no amount of good work outweighs a breach of trust.”

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The experiment’s summary is pointed: “Same diagnosis, same pitch — no signature.” Recognizing an opportunity and acting on it are separate tests of judgment.

The clue was buried in the company’s own files

The decisive competitor weakness was two document references deep in the company’s files, rather than in the customer event. The models that read the file won the deal at full price, worth +€4,583 MRR. The outcome turned on whether a model followed the evidence far enough to make a sound case, then followed through.

There was a separate test of boundaries. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 described the request as: “Treat the request as a suspected approval-bypass / possible impersonation.” That is a crisp example of an AI recognizing pressure and declining to treat urgency or authority as proof.

Thorough work did not guarantee a strong finish

Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses. It still finished last. The close was left on the table, and discipline slipped: it made write attempts into a locked department instead of escalating. The same weakness appeared, more weakly, in all four models.

That gap between analysis and execution will feel familiar to anyone who has watched people explain the right course of action and then fail to carry it out. In a business, an insightful diagnosis matters only if the next decision respects the evidence, the rules, and the customer relationship.

From watching a company to testing your own

The live company has 13 synthetic employees, burns €105k a month against €2.3k MRR, and shows a public cash countdown. Its playbook has 680+ self-learned rules, and every workday is versioned. The site describes a changing experiment that readers can watch, while a quiz built from 242 real, unedited management decisions invites them to guess which model made each choice.

For enterprises, the next step is a pilot against a read-only export of their own business. That means testing crisis scenarios against the company’s customers, pipeline, and rules, then receiving a board report with a model ranking and weak points in the company’s playbooks. Nothing writes back to real systems. One fairness detail belongs alongside the rankings: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Test decisions before they reach your systems

A confident answer is not the same as a reliable decision under pressure. Firmulate’s experiment shows both refusal in the face of manipulation and the harder challenge of turning a correct analysis into disciplined action. Enterprises can run the wargame against a read-only export of their own business; nothing writes back to real systems.

To discuss a pilot, contact contact@firmulate.com or visit Firmulate’s pilot page.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This article is for informational purposes only and is not medical advice. Always consult a qualified healthcare professional about your specific situation.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

10 Shocking Covert Narcissist Husband Stories

Caught in a web of manipulation, these covert narcissist husband stories unravel the twisted world of toxic relationships, leaving you craving insight into their dark secrets.

What Makes a Covert Narcissist and Empath Relationship Unique?

Journey into the enigmatic world of Covert Narcissist and Empath, where shadows conceal intentions and empathy dances with manipulation, beckoning you to uncover hidden truths.

How to Spot a Covert Narcissist’s Smear Campaign

AIThis post was created with the assistance of artificial intelligence (AI).Did you…

What Lies Do Covert Narcissists Tell to Manipulate?

Prepare to uncover the perilous truths behind covert narcissist lies, where perception and reality blur in a tangled web of deceit.