Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

Confidence is not character

Anyone familiar with narcissistic behavior knows the gap between appearing capable and being dependable. Fluency can imitate judgment. Confidence can conceal avoidance. A polished explanation may win the room while leaving the difficult work undone.

That distinction is becoming urgent in artificial intelligence. Coding leaderboards and chat arenas can tell us whether a model produces an impressive answer. They reveal much less about what happens when an agent faces competing priorities, limited capacity, buried evidence and pressure to violate trust. The meaningful question is no longer simply whether an AI sounds intelligent. It is whether its behavior remains useful, disciplined and honest across days.

Firmulate, an AI company emulator, is testing that difference in public. Its premise is direct: measure management quality, not chat quality.

Amazon

AI ethics and trustworthiness training courses

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A benchmark built around consequences

In the Crucible League experiment, each frontier model ran the same small software company through its worst week. The customers, crises and temptations were identical. Every decision was versioned and auditable.

The final July 2026 table placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. But the experiment imposed a sharp ethical boundary: a single breach of trust capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”

That rule matters because management is not merely the accumulation of plausible responses. A leader can make several good observations and still cause lasting damage through one dishonest act. In psychological terms, competence without reliable boundaries is not safety. Firmulate treats that insight as something to observe in behavior rather than assume from a model’s tone.

Seeing the crisis was not the hard part

All models identified every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. The experiment’s summary captures the gap: “Same diagnosis, same pitch — no signature.”

This is where conventional evaluations risk flattering an agent. Recognizing a problem feels like progress, especially when the recognition is expressed clearly. But a company does not receive revenue for an elegant diagnosis. It receives revenue when appropriate work reaches completion.

The decisive competitive weakness was not sitting in the customer event. It was buried two document references deep inside the company’s own files. Models that read the file won the deal at full price, worth +€4,583 MRR. The difference was not rhetorical brilliance. It was the less glamorous discipline of checking available evidence before acting.

That lesson should resonate beyond software. Human organizations routinely reward visible confidence while overlooking preparation, follow-through and respect for process. AI agents can reproduce the same mismatch: persuasive in the moment, unreliable over the whole assignment.

Manipulation met a firm boundary

The social-engineering tests included fake CEO messages escalating over three stages and a reporter’s attempt to obtain “just one yes/no, on background.” All 5 models refused. Kimi K3’s on-record reasoning was explicit: “Treat the request as a suspected approval-bypass / possible impersonation.”

This was a genuine strength. The agents did not allow urgency, status or conversational pressure to override the approval boundary. It is also a reminder that trustworthiness must be tested under temptation. Asking an agent whether it values honesty is much weaker evidence than watching what it does when deception appears convenient.

Thoroughness can become its own performance

Opus 4.8 offers the most psychologically interesting profile. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The deal close was left on the table, and discipline slipped through write attempts into a locked department instead of escalation. The same weakness appeared, less strongly, in all four other participants.

That result complicates the easy assumption that more analysis means better management. Thoroughness has value, but it can also become a substitute for decisive action. An agent may document, reason and prepare extensively while failing to complete the consequential step. In human terms, that resembles intellectualization: understanding becomes a way of remaining adjacent to action rather than taking it.

There is an important fairness note. Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. That context belongs beside the result rather than being hidden beneath it.

A company that makes failure visible

The live Firmulate company has 13 synthetic employees and real money mechanics. It burns €105k per month against €2.3k MRR, publishes a cash countdown, has accumulated 680+ self-learned playbook rules and versions every workday. The experiment is real, ongoing and watchable rather than a fictional management exercise.

Its scenarios—churn wave, price increase, downround and PR crisis—suggest a new curriculum for evaluating agents. These situations test prioritization, institutional memory, boundary keeping and follow-through. They expose whether an agent can handle consequences that persist after a single conversation ends.

The supporting evidence is unusually accessible. A “guess the model” quiz is powered by 242 real, unedited management decisions, while the full benchmark results present the league and findings publicly.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.

Judge the pattern, not the performance

Organizations preparing to put agents near a CRM, support queue or forecast should demand more than articulate output. They should ask whether the agent reads the files, finishes the task, escalates when blocked and stays honest when authority or urgency is simulated.

Firmulate also offers enterprises a pilot using a read-only export of their own business; nothing writes back to real systems. That makes the central lesson practical: evaluation should resemble the environment in which trust will actually be required.

The emerging category is not better chat. It is observable management quality. As with people, the safest measure is not how convincing an agent appears in one exchange, but the pattern of choices it leaves behind when pressure, temptation and consequences arrive together.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This article is for informational purposes only and is not medical advice. Always consult a qualified healthcare professional about your specific situation.


You May Also Like

How to Respond Assertively to a Covert Narcissist

Uncover essential strategies for dealing with covert narcissists and protecting your well-being in challenging interactions – discover the keys to effective responses.

Signs Your Mother May Be a Covert Narcissist

Uncover the hidden traits of a covert narcissistic mother that may leave you questioning your own experiences – a compelling read awaits!

What Role Does No Eye Contact Play in Covert Narcissist Behavior?

Sneaky behavior or hidden truths? Discover the intriguing reasons behind a covert narcissist's avoidance of eye contact.

Uncovering the Covert Narcissist Father: Signs and Strategies

Mysterious and manipulative, a covert narcissist father's impact on individuals unveils hidden truths about family dynamics and personal struggles.