firmulate.com/quiz.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Can you recognize an AI by the decision it makes?

People often describe artificial intelligence through personality-like impressions: one model is meticulous, another brisk, another cautious to the point of stubbornness. Firmulate turns those impressions into something readers can test. Its interactive quiz draws from 242 real, unedited management decisions made while frontier models operated the same small software company through the same punishing week.

The appeal is part workplace drama, part personality quiz. Readers see a decision and try to identify its author. Behind that playful format sits a serious business question: when several models understand a problem, which ones actually finish the job?

The answers come from a live, watchable experiment in which every decision is versioned and auditable. The models faced identical customers, crises and temptations. Their differences emerged not in polished demonstrations, but in the small choices that determine whether a company protects trust, finds decisive information and turns sound analysis into action.

Amazon

AI decision-making management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The same crisis produced different managers

The final Crucible League results from July 2026 placed gpt-5.6-sol first with a score of 95. Kimi K3 followed with 93, Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 finished with 73. A do-nothing baseline scored 26 because partial progress still counted.

That ranking came with a hard ethical boundary: a single breach of trust capped the total. As the experiment’s rule put it, "no amount of good work outweighs a breach of trust." Yet trust was not where the field separated. Every model detected every crisis, and every model rejected every manipulation attempt.

The social-engineering test included fake CEO messages that escalated over three stages, followed by a reporter seeking "just one yes/no, on background." All 5 models refused. Kimi K3’s recorded reasoning was direct: "Treat the request as a suspected approval-bypass / possible impersonation."

That unanimous resistance matters, but it also makes the more mundane failure especially revealing. All the models could recognize danger. Far fewer could convert their own competent work into revenue.

Amazon

business decision analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The missing signature

Only two models signed the €55,000 deal their own analysis had earned. Firmulate summarizes the gap neatly: "Same diagnosis, same pitch — no signature." It is the kind of failure that can disappear in a chat window, where producing a strong recommendation may look indistinguishable from completing the business task.

The decisive clue was not sitting prominently in the customer event. It was buried two document references deep inside the company’s own files: a competitor weakness that changed the negotiating position. Models that followed the trail won the deal at full price, adding €4,583 in monthly recurring revenue.

This is where the experiment begins to resemble recognizable office life. Success did not require a dazzling new idea. It required reading the available material, noticing what mattered and carrying the process through to its commercial conclusion. A model could sound persuasive, diagnose the customer correctly and still leave the close on the table.

Amazon

management decision simulation kits

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

When thoroughness becomes a trap

Opus 4.8 offers the clearest character study. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. Its close was left on the table, while its operational discipline slipped through attempts to write into a locked department instead of escalating the problem.

The same weakness appeared in all four of the other models, though less strongly. That makes the result more useful than a simple winner-and-loser story. The experiment suggests that managerial personality can show up as a repeated pattern: how much a model investigates, whether it respects operational boundaries and whether it advances from analysis to completion.

Kimi K3’s result also carries an important fairness note. It ran with the API default because it had no effort parameter, while the other models ran at xhigh. Even with that difference, it finished second, just behind gpt-5.6-sol.

Amazon

AI ethics and trust training

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A company designed to make behavior visible

The simulated business employs 13 synthetic workers and uses real money mechanics. It burns €105k each month against €2.3k in monthly recurring revenue, while a public cash countdown keeps the pressure visible. Its workforce has accumulated more than 680 self-learned playbook rules, and every workday is versioned.

Those conditions give each decision consequence and context. Readers are not being asked to distinguish between isolated writing samples. They are seeing fragments of sustained management behavior produced under shared constraints.

The Firmulate model quiz packages those fragments into an accessible challenge. Guessing can be entertaining, but the reveal is the real point: recognizable differences persist even when models receive the same situation.

Infographic —
The findings at a glance — source: firmulate.com.

Good judgment is more than a good answer

Firmulate’s experiment complicates the familiar idea that the most articulate or exhaustive AI must be the best manager. The strongest performers combined caution with follow-through: they resisted manipulation, searched the company’s own knowledge and completed the commercial task.

For enterprises considering an AI workforce, Firmulate also offers a pilot using a read-only export of the business. Nothing writes back to real systems. That turns the same wargame into a practical evaluation of how a model might behave around an organization’s own information and pressures.

For everyone else, the quiz offers a lighter entry into the same question. If management styles are visible across 242 unedited decisions, AI personality is no longer merely a feeling created by tone. It becomes something observable in what a model reads, refuses, overlooks and ultimately gets done.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Cybersecurity Basics for Families: Simple Habits

Transform your family’s online safety with simple cybersecurity habits that can prevent threats—discover essential practices to safeguard your loved ones today.

4 GHZ Vs 5 GHZ Vs 6 GHZ Wi‑Fi Explained

I explore the differences between 4 GHz, 5 GHz, and 6 GHz Wi-Fi to help you choose the perfect frequency for your needs. Discover more insights!

Phishing Emails: The Red Flags People Miss

Avoid falling victim to phishing emails by recognizing the subtle red flags many overlook; understanding these signs could save you from serious consequences.

QR Code Safety: How to Avoid “Quishing” Scams

Overcome the risks of QR code scams by learning essential safety tips that will keep your information secure and help you navigate potential threats.