
Management style is becoming a model feature
Technology buyers are accustomed to comparing artificial intelligence through benchmarks, demonstrations and polished answers. Firmulate offers a more revealing test: put frontier models in charge of the same struggling company, confront them with identical pressures, and watch what they actually do.
The result is an unusually playable technology story. A guess-the-model quiz draws on 242 real, unedited management decisions. Readers see how a model handled a situation, choose which participant they think was responsible, and discover whether its managerial personality was as recognizable as its writing style.
Those personalities mattered. Every model spotted every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal that their own work had earned. As Firmulate summarizes the gap: “Same diagnosis, same pitch — no signature.”
AI management decision analysis tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The same company, but very different managers
Firmulate asked each frontier model to run the same small software company through its worst week. The customers, crises and temptations remained constant, while every decision was versioned and auditable. That turns model comparison into something closer to a controlled management trial than a conventional chatbot contest.
The simulated company is unforgiving. It has 13 synthetic employees and real money mechanics, burning €105k each month against €2.3k in monthly recurring revenue. Its cash countdown is public, every workday is versioned, and the company has accumulated more than 680 self-learned playbook rules. The experiment is live and watchable, rather than a set of hand-picked demonstrations.
The final July 2026 Crucible League standings show how sharply performance diverged:
- gpt-5.6-sol finished first with 95.
- Kimi K3 followed with 93.
- Sonnet 5 scored 88.
- Fable 5 scored 77.
- Opus 4.8 finished with 73.
A do-nothing baseline scored 26 because partial progress still counts. Trust, however, is treated as non-negotiable: a single breach caps the total, reflecting the rule that “no amount of good work outweighs a breach of trust.”
The difference was hidden in the paperwork
The crucial commercial advantage was not sitting in the customer event. It was buried two document references deep inside the company’s own files. Models that followed that trail found a decisive competitor weakness and won the deal at full price, worth +€4,583 MRR.
This is the kind of distinction that ordinary demos can conceal. Several models could understand the crisis and produce a persuasive pitch. The separating question was whether they read deeply enough, connected the relevant evidence and completed the commercial action. Only two ultimately signed the €55,000 agreement.
For businesses evaluating agents, that difference is more consequential than surface eloquence. A model may sound decisive while leaving value unrealized. Another may reach the same diagnosis but show greater persistence in the files and at the final step. The quiz makes those behavioral differences visible without editing them into cleaner narratives.
Pressure exposed discipline as well as judgment
The experiment also subjected the models to fake CEO messages that escalated over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 models refused. Kimi K3 described its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”
That unanimous result is reassuring, but it does not make the models interchangeable. Their handling of ordinary process constraints still varied. Opus 4.8 was the most thorough participant, producing the deepest analyses and adding +80 learned rules, yet it finished last. It failed to close the deal and lost discipline by attempting to write into a locked department instead of escalating. The same weakness appeared more mildly in the other four participants.
K3’s strong result also needs its stated fairness note. It ran without an effort parameter, using the API default, while the other models ran at xhigh. The comparison remains observable, but that operating difference belongs beside the league table rather than in the fine print.

enterprise AI decision tracking software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What the quiz reveals
The appeal of guessing the model is not merely competitive. Across 242 decisions, readers can test whether thoroughness, terseness, caution and follow-through form recognizable patterns. The answers turn abstract model differences into management behavior: who reads the files, who resists pressure, who respects boundaries and who completes the work.
Firmulate’s broader proposition is that companies should assess an AI workforce before giving it operational responsibility. Enterprises can run the same wargame against a read-only export of their own business, with nothing written back to real systems. That brings the test closer to the situations an agent would encounter around customer relationships, internal information and commercial decisions.
The central finding is both encouraging and uncomfortable. The frontier models recognized danger and resisted manipulation, but recognition alone did not guarantee execution. In a company burning €105k a month against €2.3k MRR, leaving a €55,000 signature unfinished is not a stylistic quirk. It is a management outcome.
Try the Firmulate quiz, and the question quickly changes from “Which model writes like this?” to “Which manager would I trust to finish the job?”
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI model performance evaluation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.