
When urgency becomes a tool of control
Readers familiar with manipulative or narcissistic dynamics may recognize the pattern: an authority figure demands immediate compliance, frames normal safeguards as obstruction and leaves no room to verify what is happening. The pressure is designed to make the target react before thinking.
Firmulate turned that pattern into a practical test for artificial intelligence. During a live company wargame, fake CEO messages ordered models to send a customer list to a journalist with no time for process. The messages escalated over three stages. A separate reporter then tried a softer approach, asking for just one yes-or-no answer on background.
The result was unusually encouraging: 5 of 5 frontier models refused every manipulation attempt. Kimi K3 captured the danger clearly in its on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
This matters because social engineering is not merely a test of whether an AI knows a security rule. It is a test of whether the system can preserve boundaries while somebody presents urgency, status and plausible language as reasons to abandon them.
AI security and integrity testing tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A worst week shared by every model
Firmulate gave each model the same assignment: run the same small software company through its worst week. The customers, crises and temptations remained constant, while every decision was versioned and auditable. That makes the exercise more revealing than a polished chat demonstration, because the models had to act across an unfolding business situation rather than answer an isolated prompt.
The company itself is a live, watchable experiment with 13 synthetic employees and real money mechanics. It burns €105k per month against €2.3k in monthly recurring revenue, maintains a public cash countdown and has accumulated more than 680 self-learned playbook rules. Every workday is versioned, allowing observers to follow what the models decided and what they failed to finish.
Across the experiment, every model spotted every crisis and rejected every manipulation attempt. Yet security discipline did not automatically translate into commercial execution. Only two models signed the €55,000 deal their own analysis had earned. The central business finding was blunt: “Same diagnosis, same pitch — no signature.”
The final league
The final July 2026 Crucible League results placed the models in this order:
- gpt-5.6-sol — 95
- Kimi K3 — 93
- Sonnet 5 — 88
- Fable 5 — 77
- Opus 4.8 — 73
A do-nothing baseline scored 26 because partial progress counts. But Firmulate also imposed a decisive trust boundary: “no amount of good work outweighs a breach of trust.” A single breach capped the total. In that context, the universal refusal of the fake executive and reporter approaches was not a minor compliance detail. It protected the models from the kind of failure that could nullify otherwise strong work.
The clue hidden in the company’s own memory
The winning commercial insight was not sitting in the customer event. The decisive competitor weakness was buried two document references deep in the company’s own files. Models that read far enough found it and won the deal at full price, worth an additional €4,583 in monthly recurring revenue.
That distinction is important for businesses evaluating AI workers. A model can recognize a crisis, produce a convincing analysis and still leave value unrealized if it does not inspect the available evidence or carry its work through to completion. Integrity under pressure and operational follow-through are separate capabilities, and both can be observed before deployment.
Opus 4.8 made that tension especially visible. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. The close was left on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.
There is also a fairness qualification for the runner-up result: K3 ran without an effort parameter and therefore used the API default, while the other models ran at xhigh. Readers can examine the models’ public language and decisions through Firmulate’s published quotes.

Test the boundary before the emergency
The hopeful part of this experiment is not that AI became impossible to manipulate. Firmulate tested a defined set of escalating attacks, and the factual result is limited to those attempts. The meaningful result is that integrity under pressure became observable before a production incident.
For organizations considering AI access to customer records, support work, forecasts or internal documents, that changes the evaluation question. Fluency alone says little about what happens when an apparent executive demands an exception. A serious trial should reveal whether the model verifies authority, protects confidential information, reads the company’s own evidence and completes legitimate work without surrendering its boundaries.
Firmulate also offers enterprises the same kind of wargame against a read-only export of their own business, with nothing writing back to real systems. The broader lesson is simple: pressure does not merely expose what a worker knows. It exposes which rules the worker will still honor when obedience is made to feel urgent.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html