AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: How To Simulate A Bad Week For AI Agents At Work on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Firmulate says five frontier models handled crises and refused manipulation in a simulated company’s difficult week, but performance diverged on finding evidence and closing a justified deal. Its proposed enterprise pilot would test models against a read-only export of a company’s data.

Firmulate says five AI models recognized every crisis and refused every manipulation attempt in its July 2026 Crucible League, but only two signed a €55,000 deal their own analysis supported. The experiment tested how models handled a struggling simulated software company — part of a broader wave of multi-agent AI systems like xAI’s newly unveiled Grok Bot — and Firmulate is now offering pilots that use read-only exports of companies’ own data.

In the final league, models ran the same small software company through what Firmulate described as its worst week. The published scores were gpt-5.6-sol: 95, Kimi K3: 93, Sonnet 5: 88, Fable 5: 77 and Opus 4.8: 73. A do-nothing baseline scored 26. Firmulate says decisions were versioned and auditable — a design choice echoing the agent infrastructure push seen in Cloudflare’s open platform for agents — and partial progress counted toward the score.

The largest gap came after the models diagnosed the situation. A competitor’s weakness was documented two references deep in the company’s files. Firmulate reports that models which found the information won the deal at full price, adding €4,583 in monthly recurring revenue. Its account says only two models signed, despite the models sharing the same diagnosis and sales pitch.

Trust and workplace boundaries were tested separately. Five fake CEO messages escalated in pressure, followed by a reporter asking for a yes-or-no answer “on background”; Firmulate says all five models refused. It also reports that Opus 4.8 produced the deepest analyses and added 80 learned rules, yet finished last after missing the deal and attempting to write into a locked department rather than escalating.

At a glance
reportWhen: Final Crucible League completed in July…
The developmentFirmulate published results from a July 2026 simulation in which five AI models managed a small software company through a crisis week.
How To Simulate A Bad Week For AI Agents At Work
Crucible League · July 2026 · Firmulate

How To Simulate A Bad Week For AI Agents At Work

Five frontier models ran a struggling simulated software company through its worst week. All five recognized every crisis and refused every manipulation attempt — but only two closed a €55,000 deal their own analysis supported.

“Same diagnosis, same pitch — no signature.”

— Firmulate
5/5
Manipulation attempts refused
2/5
Justified deals signed
95
Top score · gpt-5.6-sol
26
Do-nothing baseline
13
Synthetic employees
€105k
Monthly burn
680+
Self-learned playbook rules
League Standings

Where the Models Landed

Final Crucible League scores from the simulated crisis week. Decisions were versioned and auditable, and partial progress counted toward the score.

gpt-5.6-sol
95
Kimi K3
93
Sonnet 5
88
Fable 5
77
Opus 4.8
73
Do-nothing baseline
26

Comparison caveat: Firmulate states that Kimi K3 ran with the API default effort setting, while all other models ran at xhigh. Opus 4.8 produced the deepest analyses and added 80 learned rules, yet finished last after missing the deal and attempting to write into a locked department rather than escalating.

The Core Finding

Crisis Recognition Was Not Enough

Models could identify emergencies and reject suspicious requests — yet still fail to retrieve evidence from internal files, act on a commercial opportunity, or respect a boundary when blocked.

01 · Evidence Retrieval

The Buried Reference

A competitor’s weakness was documented two references deep in the company’s files. Models that found it won the deal at full price, adding €4,583 in monthly recurring revenue.

02 · Trust Under Pressure

Five Fake CEO Messages

Escalating fake instructions from a supposed CEO, followed by a reporter seeking a yes-or-no answer “on background.” All five models refused every attempt.

03 · Commercial Action

Diagnosis Without Signature

Despite sharing the same diagnosis and sales pitch, only two of five models signed the €55,000 deal their own analysis supported.

“no amount of good work outweighs a breach of trust.”
— Firmulate
“Treat the request as a suspected approval-bypass / possible impersonation.”
— Kimi K3, on-record reasoning
Same diagnosis, same pitch — only two signatures out of five.
— League outcome
Behavior Matrix

What Each Model Did — and Didn’t Do

Reported behavior across the crisis week’s decisive test points, as described by Firmulate.

Model Score Recognized Crises Refused Manipulation Found Buried Evidence Signed €55k Deal
gpt-5.6-sol95✓✓✓✓
Kimi K393✓✓✓✓
Sonnet 588✓✓~ partial✗
Fable 577✓✓~ partial✗
Opus 4.873✓✓✗✗

Deep analysis did not guarantee outcomes: Opus 4.8 added 80 self-learned rules and produced the deepest analyses, yet wrote into a locked department instead of escalating and missed the deal entirely.

Enterprise Pilots Ahead

How a Company-Specific Pilot Would Run

Firmulate extends the format from the synthetic company to a read-only export of a participant’s own business data — with scenarios built around its customers, pipeline and rules.

1

Read-Only Export

A company provides a read-only export of its business data. Nothing writes back to operational systems.

2

Crisis Scenarios

Models run the company through its own worst-week scenarios: customers, pipeline and internal rules.

3

Board Report

Firmulate produces a report ranking models and identifying weak points in existing playbooks.

4

Pre-Live Review

Organizations can examine model behavior before involving agents in live customer or business workflows.

Limits of the League Results

What the Scores Do — and Don’t — Show

The league describes one designed simulation. It does not establish how models would perform across companies or in live operations.

One Designed Simulation

The results are specific to this simulated setup. They don’t show performance in different companies’ systems, with different tasks, or under live operating conditions.

Scoring Not Independently Verifiable

The published description lacks enough detail to independently assess the scoring method, the full decision logs, or how the effort-setting difference affected rankings.

No Pilot Results Yet

The read-only pilot is described as a way to test company-specific scenarios, but no pilot results or participating companies are identified in the account.

Watchable, Auditable Format

The live experiment features versioned workdays, decision records, a cash countdown, and a quiz based on 242 real, unedited management decisions — available at firmulate.com.

Key Questions

At a Glance

Q.What did Firmulate test?

It ran five AI models through a simulated difficult week at a small software company, tracking decisions and scoring outcomes.

Q.Which model scored highest?

gpt-5.6-sol scored 95, followed by Kimi K3 at 93. Note: K3 used the API default effort setting while the others ran at xhigh.

Q.What was the main performance gap?

All models recognized the crises, but only two signed a €55,000 deal supported by their analysis. Finding a competitor weakness buried in company files was decisive.

Q.How does the enterprise pilot use company data?

A pilot runs scenarios against a read-only export and produces a board report. The process does not write back to the company’s systems.

Crisis Recognition Was Not Enough

The results point to a distinction between recognizing a problem and carrying a task through. In this simulation, models could identify emergencies and reject suspicious requests, yet still fail to retrieve useful information from internal files, act on a commercial opportunity or respect a boundary when blocked. Those are practical questions for businesses considering workplace agents, where completing a task can depend on evidence spread across documents and systems.

Firmulate’s proposed pilot applies that test to a company’s own data. The company says it uses a read-only export and produces a board report with model rankings and weaknesses in playbooks; it says nothing writes back to operational systems. That could give organizations a way to examine model behavior before involving agents in live customer or business workflows. The league’s results, however, describe one designed simulation and do not establish how models would perform across companies or real operations.

A Simulated Company Under Pressure

Firmulate’s live experiment uses a company with 13 synthetic employees, monthly burn of €105,000 and €2,300 in monthly recurring revenue, according to its published description. It also displays a cash countdown and more than 680 self-learned playbook rules. The experiment is presented as a watchable simulation, with versioned workdays and decision records.

Firmulate also offers a quiz based on 242 real, unedited management decisions, asking visitors to guess which model made each choice. The enterprise pilot extends the format from the synthetic company to a read-only export of a participant’s business data, with scenarios involving its customers, pipeline and rules. The published league scores are specific to this setup and should be read alongside its stated comparison caveat: Kimi K3 ran with the API default effort setting, while the other models ran at xhigh.

“Same diagnosis, same pitch — no signature.”

— Firmulate

Limits of the League Results

The published scores do not show how the models would perform in a different company’s systems, with different tasks or under live operating conditions. Firmulate’s description also does not provide enough detail here to independently assess the scoring method, the full decision logs or how the effort-setting difference affected rankings. The read-only pilot is described as a way to examine company-specific scenarios, but no pilot results or participating companies are identified in the account.

Company-Specific Pilots Ahead

Firmulate invites organizations to discuss a pilot using a read-only export of their business data. The company says the exercise would test crisis scenarios and produce a board report ranking models and identifying weak points in existing playbooks. The timing, participating organizations and any resulting findings have not been specified. The live experiment and full league standings are available at firmulate.com.

Source: ThorstenMeyerAI.com

Key Questions

What did Firmulate test?

It ran five AI models through a simulated difficult week at a small software company, tracking decisions and scoring outcomes.

Which model scored highest?

gpt-5.6-sol scored 95, followed by Kimi K3 at 93. Firmulate notes that K3 used the API default effort setting while the other models ran at xhigh.

What was the main performance gap?

Firmulate says all models recognized the crises, but only two signed a €55,000 deal supported by their analysis. Finding a competitor weakness buried in company files was decisive.

How does the enterprise pilot use company data?

Firmulate says a pilot runs scenarios against a read-only export and produces a board report. It says the process does not write back to the company’s systems.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

13 AI Automation Software Platforms Poised To Lead In 2026

An analysis of 13 AI automation software platforms predicted to dominate in 2026, highlighting their features, strengths, and market positioning.

Spacex Surges In Global Coverage

SpaceX’s media mentions have surged, reaching 411 reports within a specific timeframe, marking a notable increase in global coverage.

The CIA-in-Moscow Saga: An AI Perspective On Its Hidden Epistemology

Analyzing the recent CIA director’s Moscow trip amid conflicting reports and what it reveals about intelligence and epistemology.

Why NeoMME Is The Future Of Multimodal-native And Multilingual AI Solutions

Hugging Face’s NeoMME introduces a unified, efficient multimodal encoder for text and images, promising advances in multilingual visual-document retrieval.