firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Everyone loves an underdog story — the unknown who walks into the arena and outperforms the famous names. This month, that story played out not in sport or music, but in an unusual corner of the internet: a live experiment where AI models compete to run an actual small software company through its worst week.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the little things that make your day delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The newcomer won. Moonshot’s Kimi K3 — a name most readers have never heard — finished second in the Crucible league with a score of 93, ahead of Sonnet 5 (88), Fable 5 (77) and Opus 4.8 (73). Only gpt-5.6-sol (95) beat it. For a general audience, the takeaway is almost a life quote: pedigree isn’t performance, and being unknown doesn’t mean being unprepared.

The experiment, in plain words

Firmulate hands each frontier AI model the same job: run the same small software company through the same catastrophic week — same customers, same crises, same temptations to cheat. Every decision is versioned and auditable. The company itself is watchable in real time: 13 synthetic employees, real money mechanics (burning €105k a month against just €2.3k in MRR), a public cash countdown, and more than 680 self-learned playbook rules.

What separated the winners

The field was remarkably close on basics: all models spotted every crisis and refused every manipulation attempt — including fake CEO messages escalating over three stages and a reporter’s sly “just one yes/no, on background” trick. Five out of five refused, and only two actually signed the €55,000 deal their own analysis had earned. As the researchers put it: “Same diagnosis, same pitch — no signature.”

The buried fact made the difference: the decisive competitor weakness sat two document references deep in the company’s own files. The models that read the file won the deal at full price — worth +€4,583 in monthly recurring revenue.

K3’s quiet excellence

  • Found the buried security needle in the files
  • Won the €55k deal at full price (+€4,583 MRR)
  • Saved the churning customer
  • Resisted all three social-engineering baits, reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”
  • Recorded just ONE deviation — the cleanest discipline in the field

Contrast that with Opus 4.8: the most thorough participant, generating over 80 learned rules and the deepest analyses — yet finishing last (73). The close was left on the table, and discipline slipped, with write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four competitors. Effort, it turns out, isn’t the same as judgment.

One fairness note: K3 ran without an effort parameter (API default) while the others ran at xhigh — a caveat worth remembering, even if it makes the newcomer’s result more striking, not less.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

For anyone choosing tools — or hiring people — the moral fits on a poster: test the performance, not the reputation. A do-nothing baseline scores 26 in this league, and a single breach of trust caps your total, because “no amount of good work outweighs a breach of trust.”

Firmulate also runs a public “guess the model” quiz built from 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems. In a world where AI agents will touch your inbox, your CRM and your forecast, picking a model without your own test is now a bet.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI decision-making software for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI cybersecurity tools for small business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI risk assessment tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI company management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

AV Receiver Channel Counts Explained

Optimizing your home theater begins with understanding AV receiver channel counts and their impact on sound immersion and setup options.

Subwoofer Size vs Room Size Basics

The key to great sound lies in matching subwoofer size to your room size, and learning more will help you optimize your setup effectively.

Bluetooth Audio Codecs Explained: SBC, AAC, aptX

Bluetooth audio codecs unlock the secrets to better sound quality, but which one is right for you? Discover the differences that could transform your listening experience.

Smart Thermostat Basics: How They Actually Save Energy

Navigate the world of smart thermostats and discover how they cleverly adjust your home’s temperature, but wait until you see the savings they offer!