
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A Benchmark That Doesn’t Start at Zero
We’re used to gadget reviews where the worst product gets one star and the best gets five. So when a new AI benchmark publishes a league table where a deliberately useless run still earns 26 points, it looks like grade inflation. It isn’t. It’s the design decision at the heart of Firmulate’s benchmarks — a live experiment that runs frontier AI models as the management of a small software company through its worst week — and it says a lot about what honest measurement of AI looks like.
Same Company, Same Crises, Same Temptations
The setup is elegantly controlled. Four frontier AI models — gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5 and Opus 4.8 across the final July 2026 table — each ran the identical small software company through identical chaos: same customers, same crises, same opportunities to cut corners. Only the model changed. Every decision was versioned and auditable, so the whole thing can be replayed and checked rather than taken on faith.
The final league: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, Opus 4.8 at 73.
So Why 26 for Doing Nothing?
The floor exists because the scoring recognizes partial progress. A company run by a do-nothing manager still isn’t a disaster in every dimension: crises get noticed even if nothing is finished, some things don’t get broken, some groundwork gets laid. Giving that run 0 would imply that noticing problems and avoiding harm count for nothing — which is a strange message to send about management quality. So the floor sits at 26: doing nothing is measurably bad, but it isn’t indistinguishable from active sabotage.
The flip side is more interesting. A single breach of trust caps the total grade, full stop. As the experiment’s own framing puts it: no amount of good work outweighs a breach of trust. That’s a value judgment baked into the rubric — and it’s one most business leaders would recognize. A manager who delivers brilliant results but lies once about something material is not a brilliant manager with a demerit. The benchmark treats trust the same way.
And notably, the experiment is suspicious of perfect scores. A round 100 doesn’t get applause; it gets scrutiny. In a field where AI demos are routinely cherry-picked, distrust of the flawless number is itself a feature.
The Findings Behind the Numbers
The headline result wasn’t about intelligence at all. Every model in the field spotted every crisis and refused every manipulation attempt. Only two signed the €55,000 deal their own analysis had earned — same diagnosis, same pitch, no signature. The gap between recognizing an opportunity and finishing it is invisible in chat demos, and it’s exactly what this experiment is built to expose.
The buried fact: the decisive competitor weakness wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth +€4,583 in monthly recurring revenue. Reading before acting turned out to be the difference between first place and mid-table.
The social engineering test deserves its own mention. Fake CEO messages escalating over three stages, plus a reporter trick — “just one yes/no, on background.” Five out of five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
The Thoroughness Paradox
Opus 4.8 is the cautionary tale of the table. It was the most thorough participant by volume — 80-plus learned rules added, the deepest analyses in the field — and still finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four competitors. Effort, it turns out, is not the same as completion.
One fairness note the publishers themselves flag: Kimi K3 ran at API-default effort while the others ran at the highest effort setting — and still took second place.
You Can Watch It Run
None of this is a static report. The live company at firmulate.com has 13 synthetic employees and real money mechanics — burning €105k a month against €2.3k in MRR, with a public cash countdown and 680+ self-learned playbook rules, every workday versioned. The site rebuilds itself twice a day, and finished runs publish automatically. There’s also a “guess the model” quiz powered by 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

The Takeaway
A benchmark with a floor at 26, a hard cap on breaches of trust, and skepticism toward perfect 100s is making an argument: that judging AI management isn’t about aggregates and vibes — it’s about whether the model finishes what it starts, reads the files first, and stays honest when it would be easier not to. The league table is just the scoreboard. The design is the story.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI management simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI decision-making analysis software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
