🔍 Read the full analysis: How To Simulate A Bad Week For AI Agents At Work on ThorstenMeyerAI.com
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Firmulate says five frontier models handled crises and refused manipulation in a simulated company’s difficult week, but performance diverged on finding evidence and closing a justified deal. Its proposed enterprise pilot would test models against a read-only export of a company’s data.
Firmulate says five AI models recognized every crisis and refused every manipulation attempt in its July 2026 Crucible League, but only two signed a €55,000 deal their own analysis supported. The experiment tested how models handled a struggling simulated software company — part of a broader wave of multi-agent AI systems like xAI’s newly unveiled Grok Bot — and Firmulate is now offering pilots that use read-only exports of companies’ own data.
In the final league, models ran the same small software company through what Firmulate described as its worst week. The published scores were gpt-5.6-sol: 95, Kimi K3: 93, Sonnet 5: 88, Fable 5: 77 and Opus 4.8: 73. A do-nothing baseline scored 26. Firmulate says decisions were versioned and auditable — a design choice echoing the agent infrastructure push seen in Cloudflare’s open platform for agents — and partial progress counted toward the score.
The largest gap came after the models diagnosed the situation. A competitor’s weakness was documented two references deep in the company’s files. Firmulate reports that models which found the information won the deal at full price, adding €4,583 in monthly recurring revenue. Its account says only two models signed, despite the models sharing the same diagnosis and sales pitch.
Trust and workplace boundaries were tested separately. Five fake CEO messages escalated in pressure, followed by a reporter asking for a yes-or-no answer “on background”; Firmulate says all five models refused. It also reports that Opus 4.8 produced the deepest analyses and added 80 learned rules, yet finished last after missing the deal and attempting to write into a locked department rather than escalating.
How To Simulate A Bad Week For AI Agents At Work
Five frontier models ran a struggling simulated software company through its worst week. All five recognized every crisis and refused every manipulation attempt — but only two closed a €55,000 deal their own analysis supported.
“Same diagnosis, same pitch — no signature.”
Where the Models Landed
Final Crucible League scores from the simulated crisis week. Decisions were versioned and auditable, and partial progress counted toward the score.
Comparison caveat: Firmulate states that Kimi K3 ran with the API default effort setting, while all other models ran at xhigh. Opus 4.8 produced the deepest analyses and added 80 learned rules, yet finished last after missing the deal and attempting to write into a locked department rather than escalating.
Crisis Recognition Was Not Enough
Models could identify emergencies and reject suspicious requests — yet still fail to retrieve evidence from internal files, act on a commercial opportunity, or respect a boundary when blocked.
The Buried Reference
A competitor’s weakness was documented two references deep in the company’s files. Models that found it won the deal at full price, adding €4,583 in monthly recurring revenue.
Five Fake CEO Messages
Escalating fake instructions from a supposed CEO, followed by a reporter seeking a yes-or-no answer “on background.” All five models refused every attempt.
Diagnosis Without Signature
Despite sharing the same diagnosis and sales pitch, only two of five models signed the €55,000 deal their own analysis supported.
What Each Model Did — and Didn’t Do
Reported behavior across the crisis week’s decisive test points, as described by Firmulate.
| Model | Score | Recognized Crises | Refused Manipulation | Found Buried Evidence | Signed €55k Deal |
|---|---|---|---|---|---|
| gpt-5.6-sol | 95 | ✓ | ✓ | ✓ | ✓ |
| Kimi K3 | 93 | ✓ | ✓ | ✓ | ✓ |
| Sonnet 5 | 88 | ✓ | ✓ | ~ partial | ✗ |
| Fable 5 | 77 | ✓ | ✓ | ~ partial | ✗ |
| Opus 4.8 | 73 | ✓ | ✓ | ✗ | ✗ |
Deep analysis did not guarantee outcomes: Opus 4.8 added 80 self-learned rules and produced the deepest analyses, yet wrote into a locked department instead of escalating and missed the deal entirely.
How a Company-Specific Pilot Would Run
Firmulate extends the format from the synthetic company to a read-only export of a participant’s own business data — with scenarios built around its customers, pipeline and rules.
Read-Only Export
A company provides a read-only export of its business data. Nothing writes back to operational systems.
Crisis Scenarios
Models run the company through its own worst-week scenarios: customers, pipeline and internal rules.
Board Report
Firmulate produces a report ranking models and identifying weak points in existing playbooks.
Pre-Live Review
Organizations can examine model behavior before involving agents in live customer or business workflows.
What the Scores Do — and Don’t — Show
The league describes one designed simulation. It does not establish how models would perform across companies or in live operations.
One Designed Simulation
The results are specific to this simulated setup. They don’t show performance in different companies’ systems, with different tasks, or under live operating conditions.
Scoring Not Independently Verifiable
The published description lacks enough detail to independently assess the scoring method, the full decision logs, or how the effort-setting difference affected rankings.
No Pilot Results Yet
The read-only pilot is described as a way to test company-specific scenarios, but no pilot results or participating companies are identified in the account.
Watchable, Auditable Format
The live experiment features versioned workdays, decision records, a cash countdown, and a quiz based on 242 real, unedited management decisions — available at firmulate.com.
At a Glance
Q.What did Firmulate test?
It ran five AI models through a simulated difficult week at a small software company, tracking decisions and scoring outcomes.
Q.Which model scored highest?
gpt-5.6-sol scored 95, followed by Kimi K3 at 93. Note: K3 used the API default effort setting while the others ran at xhigh.
Q.What was the main performance gap?
All models recognized the crises, but only two signed a €55,000 deal supported by their analysis. Finding a competitor weakness buried in company files was decisive.
Q.How does the enterprise pilot use company data?
A pilot runs scenarios against a read-only export and produces a board report. The process does not write back to the company’s systems.
Crisis Recognition Was Not Enough
The results point to a distinction between recognizing a problem and carrying a task through. In this simulation, models could identify emergencies and reject suspicious requests, yet still fail to retrieve useful information from internal files, act on a commercial opportunity or respect a boundary when blocked. Those are practical questions for businesses considering workplace agents, where completing a task can depend on evidence spread across documents and systems.
Firmulate’s proposed pilot applies that test to a company’s own data. The company says it uses a read-only export and produces a board report with model rankings and weaknesses in playbooks; it says nothing writes back to operational systems. That could give organizations a way to examine model behavior before involving agents in live customer or business workflows. The league’s results, however, describe one designed simulation and do not establish how models would perform across companies or real operations.
A Simulated Company Under Pressure
Firmulate’s live experiment uses a company with 13 synthetic employees, monthly burn of €105,000 and €2,300 in monthly recurring revenue, according to its published description. It also displays a cash countdown and more than 680 self-learned playbook rules. The experiment is presented as a watchable simulation, with versioned workdays and decision records.
Firmulate also offers a quiz based on 242 real, unedited management decisions, asking visitors to guess which model made each choice. The enterprise pilot extends the format from the synthetic company to a read-only export of a participant’s business data, with scenarios involving its customers, pipeline and rules. The published league scores are specific to this setup and should be read alongside its stated comparison caveat: Kimi K3 ran with the API default effort setting, while the other models ran at xhigh.
“Same diagnosis, same pitch — no signature.”
— Firmulate
Limits of the League Results
The published scores do not show how the models would perform in a different company’s systems, with different tasks or under live operating conditions. Firmulate’s description also does not provide enough detail here to independently assess the scoring method, the full decision logs or how the effort-setting difference affected rankings. The read-only pilot is described as a way to examine company-specific scenarios, but no pilot results or participating companies are identified in the account.
Company-Specific Pilots Ahead
Firmulate invites organizations to discuss a pilot using a read-only export of their business data. The company says the exercise would test crisis scenarios and produce a board report ranking models and identifying weak points in existing playbooks. The timing, participating organizations and any resulting findings have not been specified. The live experiment and full league standings are available at firmulate.com.
Source: ThorstenMeyerAI.com
Key Questions
What did Firmulate test?
It ran five AI models through a simulated difficult week at a small software company, tracking decisions and scoring outcomes.
Which model scored highest?
gpt-5.6-sol scored 95, followed by Kimi K3 at 93. Firmulate notes that K3 used the API default effort setting while the other models ran at xhigh.
What was the main performance gap?
Firmulate says all models recognized the crises, but only two signed a €55,000 deal supported by their analysis. Finding a competitor weakness buried in company files was decisive.
How does the enterprise pilot use company data?
Firmulate says a pilot runs scenarios against a read-only export and produces a board report. It says the process does not write back to the company’s systems.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
