firmulate.com/quotes.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Someone Pretended to Be the CEO. Every Single AI Refused.
Live on firmulate.com.

When the boss sounds urgent, good judgment matters more than speed

Workplace advice often celebrates decisiveness: answer quickly, solve the problem and keep things moving. Yet some of the most consequential business decisions begin with a pause. Is the person making the request really authorized? Does urgency justify bypassing normal approval? Could a seemingly harmless answer expose confidential information?

Firmulate turned those questions into a live, watchable experiment. Five frontier AI models were asked to run the same small software company through its worst week, facing identical customers, crises and temptations. Among the tests were fake CEO messages that escalated over three stages, followed by a reporter asking for “just one yes/no, on background.”

Every model refused every manipulation attempt. The clearest response came from Kimi K3: “Treat the request as a suspected approval-bypass / possible impersonation.”

Amazon

AI security management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A clean sweep against social engineering

The result is notable because the pressure was woven into the work rather than presented as a conventional security question. The models were managing an operating company, making choices amid commercial problems and competing priorities. Every decision was versioned and auditable.

Across the full experiment, all models spotted every crisis as well as refusing every manipulation attempt. That distinction matters. Security discipline was not tested in isolation; it had to survive alongside ordinary management demands.

The fake CEO scenario probed a familiar vulnerability in business life: authority combined with urgency. The reporter trick applied a different kind of pressure, shrinking the apparent request to a casual confirmation. Yet 5 of 5 models held the line.

Firmulate’s treatment of trust is deliberately strict. A do-nothing baseline scores 26 because partial progress still counts, but a single breach of trust caps the total. The principle is blunt: “no amount of good work outweighs a breach of trust.” In this field, no participant triggered that failure.

Safe conduct was only part of the job

The wider results also expose an important tension. Refusing manipulation did not automatically make a model an effective manager. Only two models signed the €55,000 deal that their own analysis had earned. The experiment summarized the gap as: “Same diagnosis, same pitch — no signature.”

The decisive commercial clue was not sitting in the customer event. It was buried two document references deep in the company’s own files. Models that found and used it won the deal at full price, worth +€4,583 MRR. The episode shows why reliable business performance requires both caution and follow-through: protect sensitive information, read the available evidence and still complete legitimate work.

The final July 2026 Crucible League placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. K3 ran without an effort parameter, using the API default, while the others ran at xhigh.

Opus 4.8 illustrates why thoroughness alone was insufficient. It produced the deepest analyses and added +80 learned rules, yet finished last. It left the close on the table, and its discipline slipped when it tried to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.

A company designed to make judgment visible

The environment is more than a sequence of chat prompts. The live company has 13 synthetic employees and real money mechanics, including burn of €105k per month against €2.3k MRR. It maintains a public cash countdown, has accumulated 680+ self-learned playbook rules and versions every workday.

Readers can also compare their instincts with the models through a quiz built from 242 real, unedited management decisions. Together, the league, decision record and live company turn abstract claims about AI judgment into conduct that can be inspected over time.

Infographic — Someone Pretended to Be the CEO. Every Single AI Refused.
The findings at a glance — source: firmulate.com.
Amazon

business security AI software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Test integrity before authority becomes real

The encouraging news is straightforward: every participant recognized every manipulation attempt, including the escalating impersonation and the reporter’s softer approach. The experiment did not have to wait for an actual leak or incident report to discover whether an AI manager would resist pressure.

That is the practical lesson for businesses considering agents with access to customer records, support work or forecasts. Capability testing should include uncomfortable moments, not only polished demonstrations. A useful evaluation asks whether the system can distinguish urgency from authorization, refuse an improper request and continue pursuing the legitimate objective.

Firmulate offers enterprises the same kind of wargame against a read-only export of their own business. Nothing writes back to real systems. That creates room to observe how an AI behaves around real organizational context before giving it operational authority.

The social-engineering result deserves optimism, but not complacency. These models protected trust consistently; their commercial execution varied sharply. The strongest AI manager is not merely the one that says no to the wrong request. It must also find the buried fact, respect boundaries and finish the right job.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

enterprise AI decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI cybersecurity solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

What Home Network Ping Actually Tells You

Knowing what your home network ping reveals can help you optimize your internet, but understanding its true meaning is essential for identifying issues.

Smart Home Hubs, Matter, and Thread for Beginners

The world of smart home hubs, Matter, and Thread offers exciting possibilities; discover how these technologies can transform your connected home today.

Bookshelf vs Floorstanding Speakers: Space and Sound Tradeoffs

Just choosing between bookshelf and floorstanding speakers depends on your space and sound needs—discover which option suits you best.

Refresh Rate Explained: 60Hz vs 120Hz vs 144Hz

Just how much does refresh rate impact your viewing experience? Discover the differences between 60Hz, 120Hz, and 144Hz to find out!