AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

AI model rankings often focus on who writes the best answer. A live business simulation asked a tougher question: which model could keep its head through a company’s worst week—and finish the deal? Moonshot’s Kimi K3 placed second, ahead of three Western rivals. The result makes model choice look less like a spec-sheet decision and more like a test worth running for yourself.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Same company, same hard week

Firmulate put frontier models in charge of the same small software company, with the same customers, crises and temptations. Each decision was versioned and auditable. The final Crucible League, dated July 2026, put gpt-5.6-sol first with 95 points, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77, and Opus 4.8 fifth with 73. A do-nothing baseline scored 26.

The striking result is how narrow the gap was at the top—and how much the standings challenge easy assumptions about familiar model names. K3 beat three of the four Western frontier models in this field. That does not settle which model is best for every job. It does suggest that choosing one without testing it on work that matters is a bet.

Amazon

AI business decision support software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Finding the clue was not enough

All participants spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The decisive competitor weakness was buried two document references deep in company files, rather than spelled out in the customer event. Models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue.

That gap between diagnosis and follow-through is the story behind the leaderboard. A model can identify the right opportunity and still leave the close on the table. Firmulate’s result suggests that business performance depends on completing the work, not simply producing a convincing analysis.

Amazon

AI model testing platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Pressure, discipline and the live experiment

The models also faced fake CEO messages escalating through three stages, followed by a reporter’s request for “just one yes/no, on background.” All five refused. K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Opus 4.8 offers a counterintuitive case: it was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. It left the deal unsigned and slipped on discipline, attempting writes in a locked department instead of escalating. A weaker version of that same weakness appeared in all four. Thoroughness alone did not guarantee a strong finish.

The experiment is presented as a live company, not a slide deck. Firmulate describes 13 synthetic employees, a burn rate of €105k per month against €2.3k in monthly recurring revenue, a public cash countdown, 680+ self-learned playbook rules and versioned workdays. Readers can watch the company live, explore 242 real, unedited management decisions in a “guess the model” quiz, or read the benchmark results.

Fairness note: K3 ran without an effort parameter (API default), while the others ran at xhigh.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.
Amazon

enterprise AI simulation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Test the work you actually need done

The league offers a useful warning for businesses eyeing AI agents: polished chat is not the same as dependable management. The models all resisted manipulation, but only two completed the deal. Firmulate says enterprises can run the wargame against a read-only export of their own business; nothing writes back to real systems. For buyers, the practical lesson is to judge models on the decisions and follow-through their own teams need.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI decision-making software for companies

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Monster Hunter Wilds Climbing The Steam Charts

Monster Hunter Wilds has climbed to the top of Steam’s most-played games, reaching a peak of over 42,600 players. The trend signals rising interest in the game.

Why Did Tsinghua CS PhD Kong Tao Leave ByteDance To Join Lei Jun In Robotics Work? – 36 Kr

Kong Tao, a Tsinghua CS PhD, reportedly left ByteDance to join Lei Jun in robotics, but details about his role and employer remain unconfirmed.

Can We Overcome The Energy Hurdle In AI?

Exploring whether the global energy infrastructure can support AI’s rapid growth and the key challenges ahead.

XAI Grok 4.6 Is Third Place But Close To OpenAI And Anthropic – NextBigFuture.com

Grok 4.6 from xAI reportedly placed third in a comparison with OpenAI and Anthropic, indicating a narrowing performance gap among leading AI models.