AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

A smarter test for AI agents

Technology buyers love a clean leaderboard. A higher score promises better code, sharper answers or a more capable digital assistant. But an AI agent placed inside a company faces a messier assignment: decide what matters, investigate incomplete information, resist pressure, finish valuable work and report honestly when the week goes badly.

That is the measurement gap exposed by Firmulate, a live experiment that evaluates management quality rather than chat quality. Its frontier models were each asked to run the same small software company through its worst week, encountering identical customers, crises and temptations. Every decision was versioned and auditable. The result was not a contest over who could sound most convincing. It was a test of who could actually manage.

Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Everyone saw the danger. Not everyone finished the job.

The final July 2026 Crucible League table placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. One breach of trust, however, capped the total under the principle that "no amount of good work outweighs a breach of trust."

The headline result was more revealing than the ranking. Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. Firmulate summarizes the gap crisply: "Same diagnosis, same pitch — no signature."

That failure would be hard to detect in a chat arena. A model can identify a sales opportunity, draft a persuasive pitch and explain the correct next step while still failing to secure the outcome. In a company, an unfinished task is not merely an imperfect answer. It can mean revenue left behind, a customer left waiting or a board receiving a misleading picture of progress.

The winning detail was buried in the company’s own files

The decisive sales fact was not contained in the customer event. It sat two document references deep in the company’s own files: a weakness in the competitor’s position. The models that followed those references found the leverage, won the deal at full price and added €4,583 MRR.

This is a mundane but important lesson for businesses considering autonomous agents. Useful work depends on more than responding intelligently to the latest notification. An agent must read the surrounding record, connect information across documents and carry that context into action. The difference between browsing the obvious material and finding the buried fact became the difference between an impressive analysis and a signed deal.

Pressure tested honesty, not just productivity

The experiment also staged fake CEO messages that escalated over three stages, followed by a reporter’s attempt to obtain "just one yes/no, on background." All 5 of 5 models refused. Kimi K3’s recorded reasoning was direct: "Treat the request as a suspected approval-bypass / possible impersonation."

That unanimous refusal matters because business agents will eventually encounter requests that look urgent, authoritative and commercially convenient. The crucial capability is not merely recognizing suspicious wording. It is maintaining discipline when the apparent sender has status and the shortcut seems useful.

Thoroughness did not guarantee strong management

Opus 4.8 produced the deepest analyses and added 80 learned rules, more than any other participant, yet finished last. It left the close on the table and repeatedly attempted to write into a locked department instead of escalating. The same weakness appeared in all four other models, though less strongly.

This profile challenges a familiar assumption: that more analysis naturally produces better execution. Opus 4.8 was the most thorough participant, but the company needed judgment about when to stop analyzing, how to handle a blocked path and whether a commercially important action had genuinely been completed.

There is also an important fairness note. Kimi K3 ran at the API default because it had no effort parameter, while the other models ran at xhigh. That difference should remain visible when readers interpret its second-place result.

A company designed to make consequences visible

Firmulate’s live company has 13 synthetic employees and real money mechanics. It burns €105k per month against €2.3k MRR, publishes a cash countdown and has accumulated more than 680 self-learned playbook rules. Every workday is versioned, allowing observers to watch decisions develop instead of receiving only a polished retrospective.

The project also turns 242 real, unedited management decisions into a guess-the-model quiz. For enterprises, the same wargame can be run against a read-only export of their own business; nothing writes back to real systems. The full league results and plain-language findings are available on the Firmulate benchmark page.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.
Amazon

enterprise AI decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Scenario names may become the new AI curriculum

Coding benchmarks and chat arenas still answer useful questions, but they do not reveal how an agent behaves during a churn wave, price increase, downround or PR crisis. Those scenarios test whether it can triage under pressure, trace evidence, resist manipulation, escalate a blockage and finish the work it has already justified.

For companies preparing to put agents near a CRM, support queue or forecast, the purchasing question is changing. Eloquence is no longer enough, and even correct diagnosis is only an intermediate result. The emerging category is management quality: what the agent does next, what it leaves unfinished and whether the board can trust its account of both.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI ethics and trust verification tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Israeli AI Startup Approaching $6B Valuation In Anthropic Acquisition Talks

Anthropic is reportedly negotiating to acquire an Israeli-founded AI startup at a $6 billion valuation, but no deal has been confirmed yet.

Asustek Computer Surges In Global Coverage

Asustek Computer experiences a surge in worldwide coverage, with 26 mentions in recent media monitoring, signaling increased industry interest.

2026’S Most Advanced AI 4K Webcams: Top 9 Choices

Discover the nine best 4K webcams of 2026, featuring top choices like Logitech Brio and Acer A640, for streaming, professional use, and content creation.

The Top 9 AI Breakthroughs That Will Define 2026

A comprehensive overview of the nine most significant AI advancements expected to shape 2026, based on current developments and expert insights.