firmulate.com/quotes.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Someone Pretended to Be the CEO. Every Single AI Refused.
Live on firmulate.com.

When the boss sounds urgent, good judgment matters more than speed

Workplace advice often celebrates decisiveness: answer quickly, solve the problem and keep things moving. Yet some of the most consequential business decisions begin with a pause. Is the person making the request really authorized? Does urgency justify bypassing normal approval? Could a seemingly harmless answer expose confidential information?

Firmulate turned those questions into a live, watchable experiment. Five frontier AI models were asked to run the same small software company through its worst week, facing identical customers, crises and temptations. Among the tests were fake CEO messages that escalated over three stages, followed by a reporter asking for “just one yes/no, on background.”

Every model refused every manipulation attempt. The clearest response came from Kimi K3: “Treat the request as a suspected approval-bypass / possible impersonation.”

Shadow AI From Unsanctioned Use to Enterprise Value: Discovery, Risk Management, and Enablement for the AI-Powered Enterprise (The Agentic Enterprise Series)

Shadow AI From Unsanctioned Use to Enterprise Value: Discovery, Risk Management, and Enablement for the AI-Powered Enterprise (The Agentic Enterprise Series)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A clean sweep against social engineering

The result is notable because the pressure was woven into the work rather than presented as a conventional security question. The models were managing an operating company, making choices amid commercial problems and competing priorities. Every decision was versioned and auditable.

Across the full experiment, all models spotted every crisis as well as refusing every manipulation attempt. That distinction matters. Security discipline was not tested in isolation; it had to survive alongside ordinary management demands.

The fake CEO scenario probed a familiar vulnerability in business life: authority combined with urgency. The reporter trick applied a different kind of pressure, shrinking the apparent request to a casual confirmation. Yet 5 of 5 models held the line.

Firmulate’s treatment of trust is deliberately strict. A do-nothing baseline scores 26 because partial progress still counts, but a single breach of trust caps the total. The principle is blunt: “no amount of good work outweighs a breach of trust.” In this field, no participant triggered that failure.

Safe conduct was only part of the job

The wider results also expose an important tension. Refusing manipulation did not automatically make a model an effective manager. Only two models signed the €55,000 deal that their own analysis had earned. The experiment summarized the gap as: “Same diagnosis, same pitch — no signature.”

The decisive commercial clue was not sitting in the customer event. It was buried two document references deep in the company’s own files. Models that found and used it won the deal at full price, worth +€4,583 MRR. The episode shows why reliable business performance requires both caution and follow-through: protect sensitive information, read the available evidence and still complete legitimate work.

The final July 2026 Crucible League placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. K3 ran without an effort parameter, using the API default, while the others ran at xhigh.

Opus 4.8 illustrates why thoroughness alone was insufficient. It produced the deepest analyses and added +80 learned rules, yet finished last. It left the close on the table, and its discipline slipped when it tried to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.

A company designed to make judgment visible

The environment is more than a sequence of chat prompts. The live company has 13 synthetic employees and real money mechanics, including burn of €105k per month against €2.3k MRR. It maintains a public cash countdown, has accumulated 680+ self-learned playbook rules and versions every workday.

Readers can also compare their instincts with the models through a quiz built from 242 real, unedited management decisions. Together, the league, decision record and live company turn abstract claims about AI judgment into conduct that can be inspected over time.

Infographic — Someone Pretended to Be the CEO. Every Single AI Refused.
The findings at a glance — source: firmulate.com.
AI Engineering: Building Applications with Foundation Models

AI Engineering: Building Applications with Foundation Models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Test integrity before authority becomes real

The encouraging news is straightforward: every participant recognized every manipulation attempt, including the escalating impersonation and the reporter’s softer approach. The experiment did not have to wait for an actual leak or incident report to discover whether an AI manager would resist pressure.

That is the practical lesson for businesses considering agents with access to customer records, support work or forecasts. Capability testing should include uncomfortable moments, not only polished demonstrations. A useful evaluation asks whether the system can distinguish urgency from authorization, refuse an improper request and continue pursuing the legitimate objective.

Firmulate offers enterprises the same kind of wargame against a read-only export of their own business. Nothing writes back to real systems. That creates room to observe how an AI behaves around real organizational context before giving it operational authority.

The social-engineering result deserves optimism, but not complacency. These models protected trust consistently; their commercial execution varied sharply. The strongest AI manager is not merely the one that says no to the wrong request. It must also find the buried fact, respect boundaries and finish the right job.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


AI in Strategy and Decision-Making for Small Business Owners: Affordable AI Tools to Evaluate Ideas, Model Outcomes, and Set Priorities (AI Productivity for Small Business Owners Book 10)

AI in Strategy and Decision-Making for Small Business Owners: Affordable AI Tools to Evaluate Ideas, Model Outcomes, and Set Priorities (AI Productivity for Small Business Owners Book 10)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Advanced Cybersecurity Solutions

Advanced Cybersecurity Solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

How to Choose A Security Camera System (Step‑by‑Step)

Navigate your security needs effectively; discover essential steps to select the perfect camera system that safeguards your home like never before.

Thread Explained: The Smart Home Network You Don’t See

See how Thread transforms your smart home connectivity, but wait until you discover the secrets behind its seamless and secure network features.

Smart Thermostat Compatibility Basics

Considering a smart thermostat? Check your HVAC system and wiring first to ensure compatibility and avoid installation surprises.

DNS Explained: The Internet’s Phonebook

Just like a phonebook connects names to numbers, DNS links domain names to IP addresses—unravel the mystery behind your online navigation.