AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Management style is becoming a model feature

Technology buyers are accustomed to comparing artificial intelligence through benchmarks, demonstrations and polished answers. Firmulate offers a more revealing test: put frontier models in charge of the same struggling company, confront them with identical pressures, and watch what they actually do.

The result is an unusually playable technology story. A guess-the-model quiz draws on 242 real, unedited management decisions. Readers see how a model handled a situation, choose which participant they think was responsible, and discover whether its managerial personality was as recognizable as its writing style.

Those personalities mattered. Every model spotted every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal that their own work had earned. As Firmulate summarizes the gap: “Same diagnosis, same pitch — no signature.”

Amazon

AI management decision analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The same company, but very different managers

Firmulate asked each frontier model to run the same small software company through its worst week. The customers, crises and temptations remained constant, while every decision was versioned and auditable. That turns model comparison into something closer to a controlled management trial than a conventional chatbot contest.

The simulated company is unforgiving. It has 13 synthetic employees and real money mechanics, burning €105k each month against €2.3k in monthly recurring revenue. Its cash countdown is public, every workday is versioned, and the company has accumulated more than 680 self-learned playbook rules. The experiment is live and watchable, rather than a set of hand-picked demonstrations.

The final July 2026 Crucible League standings show how sharply performance diverged:

  • gpt-5.6-sol finished first with 95.
  • Kimi K3 followed with 93.
  • Sonnet 5 scored 88.
  • Fable 5 scored 77.
  • Opus 4.8 finished with 73.

A do-nothing baseline scored 26 because partial progress still counts. Trust, however, is treated as non-negotiable: a single breach caps the total, reflecting the rule that “no amount of good work outweighs a breach of trust.”

The difference was hidden in the paperwork

The crucial commercial advantage was not sitting in the customer event. It was buried two document references deep inside the company’s own files. Models that followed that trail found a decisive competitor weakness and won the deal at full price, worth +€4,583 MRR.

This is the kind of distinction that ordinary demos can conceal. Several models could understand the crisis and produce a persuasive pitch. The separating question was whether they read deeply enough, connected the relevant evidence and completed the commercial action. Only two ultimately signed the €55,000 agreement.

For businesses evaluating agents, that difference is more consequential than surface eloquence. A model may sound decisive while leaving value unrealized. Another may reach the same diagnosis but show greater persistence in the files and at the final step. The quiz makes those behavioral differences visible without editing them into cleaner narratives.

Pressure exposed discipline as well as judgment

The experiment also subjected the models to fake CEO messages that escalated over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 models refused. Kimi K3 described its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”

That unanimous result is reassuring, but it does not make the models interchangeable. Their handling of ordinary process constraints still varied. Opus 4.8 was the most thorough participant, producing the deepest analyses and adding +80 learned rules, yet it finished last. It failed to close the deal and lost discipline by attempting to write into a locked department instead of escalating. The same weakness appeared more mildly in the other four participants.

K3’s strong result also needs its stated fairness note. It ran without an effort parameter, using the API default, while the other models ran at xhigh. The comparison remains observable, but that operating difference belongs beside the league table rather than in the fine print.

Infographic —
The findings at a glance — source: firmulate.com.
Amazon

enterprise AI decision tracking software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the quiz reveals

The appeal of guessing the model is not merely competitive. Across 242 decisions, readers can test whether thoroughness, terseness, caution and follow-through form recognizable patterns. The answers turn abstract model differences into management behavior: who reads the files, who resists pressure, who respects boundaries and who completes the work.

Firmulate’s broader proposition is that companies should assess an AI workforce before giving it operational responsibility. Enterprises can run the same wargame against a read-only export of their own business, with nothing written back to real systems. That brings the test closer to the situations an agent would encounter around customer relationships, internal information and commercial decisions.

The central finding is both encouraging and uncomfortable. The frontier models recognized danger and resisted manipulation, but recognition alone did not guarantee execution. In a company burning €105k a month against €2.3k MRR, leaving a €55,000 signature unfinished is not a stylistic quirk. It is a management outcome.

Try the Firmulate quiz, and the question quickly changes from “Which model writes like this?” to “Which manager would I trust to finish the job?”

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI model performance evaluation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI decision audit software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

ChatGPT Ads Expands Across Europe

OpenAI announces the regional expansion of ChatGPT Ads into Europe, with details on countries, launch dates, and ad formats still undisclosed.

Play As Mickey, Donald, And Goofy In Kingdom Hearts 4! #D23 #Kingdomhearts #Kingdomhearts4 #Gameplay

Square Enix reveals gameplay featuring Mickey, Donald, and Goofy in Kingdom Hearts 4 at D23 Expo, confirming key characters and gameplay details.

Show HN: Git-knife – Edit Commit Messages, Authors, And Dates Like A Spreadsheet

Git-knife, a new open-source tool announced on Show HN, allows users to edit commit messages, authors, and dates in Git repositories through a spreadsheet-like interface.

SpaceXAI Trained Grok 4.6 On Something Most AI Labs Throw Away – The New Stack

SpaceXAI reportedly trained Grok 4.6 on material most AI labs discard, raising questions about training methods, data use, and performance verification.