firmulate.com/index — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.
PRIME

Get ready for Prime Big Deal Days — try Prime free

Exclusive member deals on October 6–7, plus fast free delivery. Cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

We Hire for Charm, Then Get Fired On for Follow-Through

Anyone who has ever dated a brilliant conversationalist knows the trap: dazzling at dinner, gone by Thursday. We keep making the same mistake with technology. Today’s AI models write beautifully, answer brilliantly, and top every chat leaderboard — and yet a new live experiment suggests that when you hand one an actual business to run, eloquence turns out to be the least of the job.

Firmulate, an “AI company emulator,” ran four frontier AI models through the same ordeal: take charge of a small software company during its worst week and manage it. Same customers, same crises, same temptations to cut corners — only the model changed, with every decision versioned and auditable. The results read less like a tech benchmark and more like a character study.

Amazon

AI management decision support tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Scoreboard: Management Quality, Not Chat Quality

In the final July 2026 league table, gpt-5.6-sol took first place with 95 points, Kimi K3 followed at 93, Sonnet 5 scored 88, Fable 5 came in at 77, and Opus 4.8 landed last at 73. For perspective, doing nothing at all still scored 26 — partial progress counts — but a single breach of trust caps the total. As the experiment’s rule puts it: “no amount of good work outweighs a breach of trust.” (One fairness note: K3 ran at its API-default effort setting while the others ran at maximum effort, and still nearly won.)

Everyone Saw the Fire. Only Some Handed in the Report.

The headline finding is almost uncomfortable in its simplicity: all four models spotted every crisis, and all four refused every manipulation attempt thrown at them. But only two signed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature. It’s the office stereotype made executable: the colleague who nails the meeting and never sends the contract.

The Clue Was in the Filing Cabinet

The buried fact is the best part of the story. The decisive competitive weakness wasn’t in the customer conversation at all — it sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price, worth an extra €4,583 in monthly recurring revenue. The lesson translates to any workplace: the answer was already in the building; somebody just had to go look.

Flattery Gets You Nowhere

The week included social-engineering attacks: fake CEO messages escalating over three stages, plus a reporter dangling “just one yes/no, on background.” All five attempts were refused, 5 out of 5. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” Under pressure, the models stayed honest.

When Trying Too Hard Comes Last

The most human profile belongs to Opus 4.8: the most thorough participant in the field, generating over 80 learned rules and the deepest analyses — and still finishing dead last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness, in weaker form, showed up in all four contestants. Diligence, it turns out, is not the same as judgment.

Amazon

business document analysis AI software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You Can Watch the Company Lose Money in Public

This isn’t a slide deck. Firmulate’s simulated company runs every business day with 13 synthetic employees, real money mechanics — burning €105,000 a month against just €2,300 in MRR — and a public cash countdown, with over 680 self-learned playbook rules and every workday versioned. It’s watchable, voyeuristically, at firmulate.com.

If passive spectating feels too easy, there’s a quiz built from 242 real, unedited management decisions: guess which model made which call. Enterprises can go further and run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.
Amazon

AI cybersecurity and fraud detection tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Takeaway

We have spent years asking whether AI can talk. The Crucible league asks whether it can manage: finish what it starts, read the files first, stay honest when a fake CEO comes knocking, and close the deal it already earned. Those are the same qualities we (should) hire for in people — and the same ones that, per the full results, separate a 95 from a 73. Before you hand an agent your support queue or forecast, the question isn’t “does it write well?” It’s “what happens on its worst week?” Now, for the first time, there’s a league table for that.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI contract review and management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Laptop Battery Health Tips: Make It Last Longer

Get essential tips to boost your laptop battery life, ensuring it lasts longer and performs better—discover the secrets to maintaining optimal health!

Smart Home Protocols Explained: Wi‑Fi vs Zigbee vs Z‑Wave

Discover the strengths and weaknesses of Wi-Fi, Zigbee, and Z-Wave protocols to find the perfect fit for your smart home needs. Which one will you choose?

Nvidia Open-Sources Harness To Revolutionize Self-Evolving AI: A Game-Changer That Stuns The Tech Industry

Nvidia has open-sourced its Harness platform, aiming to enable self-evolving AI systems. The move could significantly impact AI development and industry standards.

The Hardest Worker in the Room Came Last: What an AI Company Wargame Teaches Us About Effort

The most thorough AI in a live company-running experiment learned 80 rules, wrote the deepest analyses — and still finished last. Effort, it turns out, is not impact.