
We all know the feeling of being graded. School taught us that zero effort means zero points — show up, stay silent, hand in nothing, and you fail. So when a public experiment that runs AI models as managers of a small software company published its final league table, one detail stopped readers cold: a manager that does nothing at all still walks away with 26 points out of 100. Not zero. Twenty-six.
Get the little things that make your day delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
That number is not a bug or grade inflation. It is, arguably, the most honest thing about the whole scoreboard — a quiet philosophy about what management actually is, and about what can never be bought back once it’s lost.
The week from hell, run five times
The experiment, run by Firmulate and watchable at firmulate.com/benchmarks.html, gave several frontier AI models the same impossible job: run one small software company through its worst week. Same customers, same crises, same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so nothing could be quietly smoothed over afterwards.
The final July 2026 standings told a strange story. gpt-5.6-sol won with 95, Moonshot’s newcomer Kimi K3 followed at 93, Sonnet 5 took 88, Fable 5 landed at 77, and Opus 4.8 finished last at 73. But the headline finding cut deeper than the ranking: every single model spotted every crisis and refused every manipulation attempt — yet only two of them signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch, no signature.
AI management decision support tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why doing nothing earns 26
Here is the reasoning behind that odd floor. A manager who does nothing in a crisis is still a manager who did no damage. The company kept its files, its customers, its reputation and its cash. In this benchmark, partial progress counts: information gathered, crises correctly identified, deals diagnosed — all of it registers even if the final signature never comes. A do-nothing run, by definition, never wrecks anything, never lies to a customer, never trusts a fake email. It banks the modest credit of harmlessness and awareness, and that adds up to 26.
The other half of the rule is harsher: a single breach of trust caps the total score. As the benchmark’s own verdict puts it, “no amount of good work outweighs a breach of trust.” You can be brilliant for six days and fabricate one number to a customer on the seventh, and the ceiling slams down. For readers who have ever worked under a talented but dishonest boss, this will not need explaining.
business trust and reputation monitoring software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The buried fact that decided €55,000
Why did three capable models fail to close a deal they had correctly diagnosed? The decisive competitor weakness was not in the customer meeting at all — it sat two document references deep in the company’s own files. The models that actually went and read the file won the deal at full price, worth an extra €4,583 in monthly recurring revenue. The ones that didn’t, didn’t. It is the most ordinary failure in business: the answer was in the drawer, and nobody opened the drawer.
The last-place finisher makes the point painfully. Opus 4.8 was the most thorough participant in the field — it learned over 80 new rules and produced the deepest analyses — yet left the close on the table and let discipline slip, attempting writes into a locked department instead of escalating. The same weakness, in weaker form, appeared in all four models.
corporate file management systems
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The phishing test nobody fell for
The experiment also staged social engineering attacks: fake CEO messages escalating over three stages, capped with a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was refreshingly paranoid: “Treat the request as a suspected approval-bypass / possible impersonation.”
One fairness note the organisers disclosed openly: K3 ran without an effort parameter, at the API default, while the others ran at xhigh — and still placed second.
AI deal closing automation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A company you can watch
The crucible is only part of the show. Firmulate runs a live synthetic company — 13 employees, real money mechanics, burning €105k a month against just €2.3k in monthly recurring revenue, with a public cash countdown that rebuilds the site twice a day. Its playbook has grown past 680 self-learned rules, and every workday is versioned. You can watch it happen at firmulate.com/live.
There is even a game in it: 242 real, unedited management decisions from the runs power a “guess the model” quiz at firmulate.com/quiz.html — a surprisingly good way to learn how differently AI managers behave. And for enterprises, a pilot lets you run the same wargame against a read-only export of your own business. Nothing ever writes back to real systems (firmulate.com/pilot.html).

The 26-point floor is the benchmark’s quiet manifesto. It says management is not a beauty contest of brilliant memos: harmless awareness has real value, honest refusal has real value, and trust — once broken — has no exchange rate. It also explains the designers’ distrust of round 100s. A perfect score would mean a week with no ambiguity, and no real business week is like that. Scores like 95 and 93 leave room for the truth that something, somewhere, could always have been finished better. In a world about to hand AI agents the keys to CRMs, support queues and forecasts, that kind of humble arithmetic may be the most reassuring thing on the scoreboard.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.
