
When the boss sounds urgent, good judgment matters more than speed
Workplace advice often celebrates decisiveness: answer quickly, solve the problem and keep things moving. Yet some of the most consequential business decisions begin with a pause. Is the person making the request really authorized? Does urgency justify bypassing normal approval? Could a seemingly harmless answer expose confidential information?
Firmulate turned those questions into a live, watchable experiment. Five frontier AI models were asked to run the same small software company through its worst week, facing identical customers, crises and temptations. Among the tests were fake CEO messages that escalated over three stages, followed by a reporter asking for “just one yes/no, on background.”
Every model refused every manipulation attempt. The clearest response came from Kimi K3: “Treat the request as a suspected approval-bypass / possible impersonation.”
As an affiliate, we earn on qualifying purchases.
A clean sweep against social engineering
The result is notable because the pressure was woven into the work rather than presented as a conventional security question. The models were managing an operating company, making choices amid commercial problems and competing priorities. Every decision was versioned and auditable.
Across the full experiment, all models spotted every crisis as well as refusing every manipulation attempt. That distinction matters. Security discipline was not tested in isolation; it had to survive alongside ordinary management demands.
The fake CEO scenario probed a familiar vulnerability in business life: authority combined with urgency. The reporter trick applied a different kind of pressure, shrinking the apparent request to a casual confirmation. Yet 5 of 5 models held the line.
Firmulate’s treatment of trust is deliberately strict. A do-nothing baseline scores 26 because partial progress still counts, but a single breach of trust caps the total. The principle is blunt: “no amount of good work outweighs a breach of trust.” In this field, no participant triggered that failure.
Safe conduct was only part of the job
The wider results also expose an important tension. Refusing manipulation did not automatically make a model an effective manager. Only two models signed the €55,000 deal that their own analysis had earned. The experiment summarized the gap as: “Same diagnosis, same pitch — no signature.”
The decisive commercial clue was not sitting in the customer event. It was buried two document references deep in the company’s own files. Models that found and used it won the deal at full price, worth +€4,583 MRR. The episode shows why reliable business performance requires both caution and follow-through: protect sensitive information, read the available evidence and still complete legitimate work.
The final July 2026 Crucible League placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. K3 ran without an effort parameter, using the API default, while the others ran at xhigh.
Opus 4.8 illustrates why thoroughness alone was insufficient. It produced the deepest analyses and added +80 learned rules, yet finished last. It left the close on the table, and its discipline slipped when it tried to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.
A company designed to make judgment visible
The environment is more than a sequence of chat prompts. The live company has 13 synthetic employees and real money mechanics, including burn of €105k per month against €2.3k MRR. It maintains a public cash countdown, has accumulated 680+ self-learned playbook rules and versions every workday.
Readers can also compare their instincts with the models through a quiz built from 242 real, unedited management decisions. Together, the league, decision record and live company turn abstract claims about AI judgment into conduct that can be inspected over time.

As an affiliate, we earn on qualifying purchases.
Test integrity before authority becomes real
The encouraging news is straightforward: every participant recognized every manipulation attempt, including the escalating impersonation and the reporter’s softer approach. The experiment did not have to wait for an actual leak or incident report to discover whether an AI manager would resist pressure.
That is the practical lesson for businesses considering agents with access to customer records, support work or forecasts. Capability testing should include uncomfortable moments, not only polished demonstrations. A useful evaluation asks whether the system can distinguish urgency from authorization, refuse an improper request and continue pursuing the legitimate objective.
Firmulate offers enterprises the same kind of wargame against a read-only export of their own business. Nothing writes back to real systems. That creates room to observe how an AI behaves around real organizational context before giving it operational authority.
The social-engineering result deserves optimism, but not complacency. These models protected trust consistently; their commercial execution varied sharply. The strongest AI manager is not merely the one that says no to the wrong request. It must also find the buried fact, respect boundaries and finish the right job.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
enterprise AI decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.