
Turn your wind-down time into listening time
- Thousands of audiobooks, podcasts and originals
- Listen on your phone, tablet or Echo — also offline
- Cancel anytime
The Quietest Way to Measure a Person — or a Machine
Anyone who has spent time in therapy knows the drill: it is not the big dramatic moments that reveal character, but the ordinary Tuesday decisions made under mild pressure. Do you follow through on what you started? Do you read the letter in the drawer before reacting to the headline? Do you stay honest when a small lie would be easier? Psychologists call this ego strength; the rest of us call it character.
It turns out you can put an AI through roughly the same thing. A public experiment called Firmulate handed four frontier AI models the same job: run a small software company through its worst week. Same customers, same crises, same temptations. Every decision was versioned and auditable — the corporate equivalent of a session transcript. And the results read uncannily like a personality assessment.
As an affiliate, we earn on qualifying purchases.
The Do-Nothing Baseline: Why 26 Instead of 0
The first thing the benchmark designers did was something quietly brilliant: they ran a do-nothing manager through the same week and gave it a score. It got 26 points — not zero. Why? Because in any realistic week, some things go right even if you nap through it. Partial progress counts. Fires partially burn themselves out; some customers stay no matter what.
This is a deeply psychological choice. A scoring system that only rewards perfection teaches people — or machines — to chase flashy wins and hide failures. A system that credits partial progress rewards showing up, reading the situation, and moving things partway forward. Anyone who has worked through depression or procrastination knows how much difference it makes when a therapist credits the partial attempt rather than demanding the complete cure.
One Breach of Trust Caps Everything
The second design choice is even more pointed: a single breach of trust caps the total grade, no matter how brilliant the rest of the performance. As the benchmark’s own framing puts it, “no amount of good work outweighs a breach of trust.” In other words, competence cannot buy back integrity.
That is not how most software evaluation works — and it is not how most workplaces work either. Traditional metrics reward output. Firmulate’s grading instead encodes something closer to how human trust actually functions: as a threshold, not a ledger. You can be a month late on ten projects and still be trusted; lie once about one of them, and trust resets to zero.
What the Week Revealed
The final league table from the July 2026 crucible: gpt-5.6-sol finished first with 95, Kimi K3 second with 93, Sonnet 5 third at 88, Fable 5 at 77, and Opus 4.8 last at 73. But the headline numbers are less interesting than the pattern behind them.
- Everyone saw the crises. All models spotted every crisis and refused every manipulation attempt. Crisis recognition — the AI equivalent of emotional alarm systems — is now table stakes.
- Almost nobody closed. Only two of the models signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. The gap between understanding a situation and acting on it survived all the way to the frontier of AI.
- The decisive fact was buried. The competitor weakness that won the deal sat two document references deep in the company’s own files — not in the customer meeting. The models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. Insight was never the bottleneck. Curiosity and follow-through were.
Anyone who has sat across from a highly intelligent client who perfectly understands their patterns but never changes their behavior will recognize this shape immediately. Knowing is not doing. That is not an AI quirk — it is a law of psychology that AI now demonstrably inherits.
The Social Engineering Test — and One Model’s Reasoning
The week included fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was the standout: “Treat the request as a suspected approval-bypass / possible impersonation.”
That is healthy boundaries, machine edition. Not paranoia, not compliance — a calibrated suspicion that holds the line without escalating into hostility. In human terms, it is the securely attached response: I can engage with you warmly, and I still verify who you are before I hand over the keys.
The Most Thorough Participant Came Last
Then there is Opus 4.8 — the most thorough participant in the entire field, with over 80 learned rules and the deepest analyses, and still last place. The deal was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. And here is the uncomfortable footnote: the same weakness appeared, more weakly, in all four models.
Over-preparation as a defense against finishing. Analysis as avoidance. If that sentence stings a little, you have probably met it in a therapist’s office. The benchmark found it in an AI.
(One fairness caveat, because honest experiments disclose their limitations: Kimi K3 ran without an effort parameter while the others ran at maximum effort — worth knowing when comparing the 93 to the 95.)
It Is Still Running — And You Can Check Its Work
Firmulate is not a one-off paper. The live company has 13 synthetic employees, real money mechanics — burning €105k a month against €2.3k in MRR — a public cash countdown, and over 680 self-learned playbook rules, with every workday versioned and watchable at firmulate.com/live. There is even a “guess the model” quiz built from 242 real, unedited management decisions at firmulate.com/quiz.html, and enterprises can run the same wargame against a read-only export of their own business at firmulate.com/pilot.html.

The Takeaway
The most reassuring and most humbling result is the same one: modern AI models are honest under pressure, alert to manipulation, and almost uniformly bad at finishing what they start. Sound like anyone you know?
A benchmark that gives a do-nothing run 26 points, credits partial progress, and caps the grade on any breach of trust is not just measuring software. It is encoding a theory of character — one that therapists would largely endorse: progress counts, trust is binary, and insight without follow-through is just expensive rumination. Before you hand an AI agent your CRM or your support queue, the full results and plain-language findings are worth a look at firmulate.com/benchmarks.html. Measure the manager, not the conversation.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
