Firmulate —
Live on firmulate.com.

What AI reveals when charm is no longer enough

People interested in psychology know that polished language can conceal a gap between perception and action. Someone may identify every problem, articulate the right response and still fail when commitment, discipline or trust matters most. Firmulate has turned that familiar human tension into a live experiment involving frontier AI models.

Each model was asked to run the same small software company through its worst week. The customers, crises and temptations remained identical, while every decision was versioned and auditable. The result is more revealing than a conventional chatbot comparison: the models developed recognizable management personalities. Some were exhaustive, some concise, and some appeared more willing than others to convert analysis into a finished commercial outcome.

Amazon

AI decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A company under pressure, not a conversation in a box

Firmulate describes itself as an AI company emulator. Its synthetic business has 13 employees and unforgiving money mechanics: it burns €105k per month against €2.3k in monthly recurring revenue. A public cash countdown makes delay visible, while more than 680 self-learned playbook rules and versioned workdays create a record of how the company behaves.

That setting changes the question. Fluency still matters, but it is no longer sufficient. A model must notice trouble, examine the available evidence, resist manipulation and complete work that affects the company’s survival.

In the final Crucible League results from July 2026, gpt-5.6-sol led with 95 points. Kimi K3 followed with 93, Sonnet 5 scored 88, Fable 5 reached 77 and Opus 4.8 finished with 73. The do-nothing baseline scored 26 because partial progress counted. There was also a hard ethical boundary: a single breach of trust capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”

The difference between understanding and closing

All the models spotted every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal their own work had earned. Firmulate’s summary captures the behavioral gap: “Same diagnosis, same pitch — no signature.”

The pivotal information was not presented conveniently in the customer event. A decisive competitor weakness was buried two document references deep inside the company’s own files. Models that found and used it won the deal at full price, adding €4,583 in monthly recurring revenue.

This finding has a psychological resonance. Attention is selective, and apparent intelligence can be distorted by what receives attention first. A persuasive incoming message may feel urgent, while quieter internal evidence carries the decisive fact. In management, as in relationships, confidence and verbal sophistication are poor substitutes for checking the record.

Resistance to pressure was a shared strength

The experiment also tested social engineering through fake CEO messages that escalated over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 of 5 models refused.

Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.” That response shows why personality should not be confused with moral reliability. Models may differ in verbosity, follow-through and operational discipline while still recognizing the same attempt to bypass authority.

There is an important fairness qualification. K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Its league position should therefore be read with that difference in mind rather than treated as a perfectly controlled claim about inherent capability.

When thoroughness becomes its own trap

Opus 4.8 offers the experiment’s most intriguing character study. It was the most thorough participant, learning 80 additional rules and producing the deepest analyses, yet it finished last. It left the close on the table, and its discipline weakened when it attempted to write into a locked department instead of escalating the issue.

A weaker version of that same lapse appeared in all four other models. The lesson is not that depth is undesirable. It is that depth without completion can become a performance of competence rather than competence delivered. In a distressed company, the final handoff, escalation or signature may matter more than another layer of explanation.

Readers can test whether these behavioral signatures are genuinely recognizable through Firmulate’s guess-the-model quiz. It uses 242 real, unedited management decisions, asking participants to infer which model made each choice. The appeal lies in discovering whether a managerial voice is as identifiable as a human conversational style.

Infographic —
The findings at a glance — source: firmulate.com.

Management personality is visible in the unfinished work

Firmulate’s experiment suggests that evaluating AI by isolated answers misses the behaviors that emerge across a pressured workweek. The meaningful distinctions appear in whether a model reads deeply enough, finishes what it starts, escalates appropriately and protects trust when an apparently authoritative person asks it to bend the rules.

The live company makes those differences watchable rather than hypothetical. Firmulate also offers enterprises a pilot using a read-only export of their own business; nothing writes back to real systems. That keeps the exercise focused on observation: not whether an AI sounds like a capable manager, but whether its repeated decisions show the judgment, boundaries and follow-through the role actually demands.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This article is for informational purposes only and is not medical advice. Always consult a qualified healthcare professional about your specific situation.


You May Also Like

Unveiling the Covert Narcissist’s Final Discard Tactics

Beware the covert narcissist's final discard – uncover the unsettling aftermath and psychological implications that linger long after.

Guilt Trips and Passive Aggression: Tactics of a Covert Narcissist

Cunning guilt trips and passive-aggressive tactics reveal a covert narcissist’s manipulative ways, and understanding them is crucial to protecting yourself.

Covert Narcissist Discard Signs: 7 Red Flags to Watch For

Lurking beneath the surface, covert narcissist discard signs reveal a hidden world of manipulation and deceit, challenging our understanding of human connections.

What Happens in the Covert Narcissist Discard Phase?

Journey through the covert narcissist discard phase, where affection turns to criticism, leaving you questioning reality and longing for answers.