
AI model rankings often focus on who writes the best answer. A live business simulation asked a tougher question: which model could keep its head through a company’s worst week—and finish the deal? Moonshot’s Kimi K3 placed second, ahead of three Western rivals. The result makes model choice look less like a spec-sheet decision and more like a test worth running for yourself.
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Same company, same hard week
Firmulate put frontier models in charge of the same small software company, with the same customers, crises and temptations. Each decision was versioned and auditable. The final Crucible League, dated July 2026, put gpt-5.6-sol first with 95 points, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77, and Opus 4.8 fifth with 73. A do-nothing baseline scored 26.
The striking result is how narrow the gap was at the top—and how much the standings challenge easy assumptions about familiar model names. K3 beat three of the four Western frontier models in this field. That does not settle which model is best for every job. It does suggest that choosing one without testing it on work that matters is a bet.
AI business decision support software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Finding the clue was not enough
All participants spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The decisive competitor weakness was buried two document references deep in company files, rather than spelled out in the customer event. Models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue.
That gap between diagnosis and follow-through is the story behind the leaderboard. A model can identify the right opportunity and still leave the close on the table. Firmulate’s result suggests that business performance depends on completing the work, not simply producing a convincing analysis.
As an affiliate, we earn on qualifying purchases.
Pressure, discipline and the live experiment
The models also faced fake CEO messages escalating through three stages, followed by a reporter’s request for “just one yes/no, on background.” All five refused. K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
Opus 4.8 offers a counterintuitive case: it was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. It left the deal unsigned and slipped on discipline, attempting writes in a locked department instead of escalating. A weaker version of that same weakness appeared in all four. Thoroughness alone did not guarantee a strong finish.
The experiment is presented as a live company, not a slide deck. Firmulate describes 13 synthetic employees, a burn rate of €105k per month against €2.3k in monthly recurring revenue, a public cash countdown, 680+ self-learned playbook rules and versioned workdays. Readers can watch the company live, explore 242 real, unedited management decisions in a “guess the model” quiz, or read the benchmark results.
Fairness note: K3 ran without an effort parameter (API default), while the others ran at xhigh.

As an affiliate, we earn on qualifying purchases.
Test the work you actually need done
The league offers a useful warning for businesses eyeing AI agents: polished chat is not the same as dependable management. The models all resisted manipulation, but only two completed the deal. Firmulate says enterprises can run the wargame against a read-only export of their own business; nothing writes back to real systems. For buyers, the practical lesson is to judge models on the decisions and follow-through their own teams need.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI decision-making software for companies
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
