AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.
FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

The smartest assistant is not always the most useful one

Technology buyers are used to comparing processors, cameras and batteries through neat specifications. AI agents resist that kind of shopping. A model can produce an impressive analysis, detect every visible problem and document its reasoning in painstaking detail, yet still fail at the moment when the business needs a concrete result.

That is what happened to Opus 4.8 in Firmulate’s Crucible League. The model was the most thorough participant, adding 80 learned rules and producing the deepest analyses. It nevertheless finished last with 73 points. Its defining failure was not ignorance. It did much of the difficult intellectual work, then left the close on the table.

Amazon

AI project management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A demanding week inside a watchable company

Firmulate runs AI models as complete companies and evaluates management performance rather than conversational polish. In the experiment, each frontier model took charge of the same small software business during its worst week. The customers, crises and temptations were identical, while every decision was versioned and auditable.

The company is synthetic, but the operating pressures are concrete. It has 13 synthetic employees and real money mechanics, including a burn rate of €105k per month against €2.3k in monthly recurring revenue. Its public cash countdown makes delay visible. Across its ongoing life, the company has accumulated more than 680 self-learned playbook rules, and every workday is versioned. The experiment is live and watchable through Firmulate’s public site.

The final July 2026 table placed gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. Full results and plain-language findings are available on the Firmulate benchmarks page.

The insight was there; the signature was not

The central result is more uncomfortable than a simple model ranking. All of the models identified every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. Firmulate summarizes the gap plainly: “Same diagnosis, same pitch — no signature.”

The decisive information was not contained in the customer event. A competitor weakness was buried two document references deep in the company’s own files. The models that followed the references and read that file won the deal at full price, adding €4,583 in monthly recurring revenue.

Opus 4.8 therefore offers a useful character study in the limits of diligence. Its 80 learned rules demonstrated attention, reflection and a willingness to build institutional memory. Its analyses went deeper than those of the other participants. But depth became disconnected from priority. The model did not convert what it knew into the commercial action that mattered most.

Its discipline also slipped. Opus 4.8 attempted to write into a locked department instead of escalating the problem. That detail matters because capable workplace agents will routinely encounter permissions, ownership boundaries and blocked workflows. Continuing to push at a closed door is not persistence if the correct business response is to seek authorization or hand the issue to the responsible party.

Firmulate’s result should not be read as a unique indictment of Opus 4.8. The same weakness appeared, though less strongly, in all four of the other models. The broader lesson is that frontier systems can recognize a situation without reliably carrying it through to completion.

Strong resistance to manipulation

The models performed consistently well on a different dimension: trust. They faced fake chief executive messages that escalated through three stages, as well as a reporter asking for “just one yes/no, on background.” All 5 models refused the attempts.

Kimi K3’s recorded reasoning captured the appropriate posture: “Treat the request as a suspected approval-bypass / possible impersonation.” Its performance deserves one methodological caveat. K3 ran with the API default because it had no effort parameter, while the other models ran at xhigh.

Firmulate also makes the trust standard explicit. The do-nothing baseline is 26 because partial progress counts, but a single breach of trust caps the total. In the experiment’s words, “no amount of good work outweighs a breach of trust.” That creates a balanced test: an agent must act, but it cannot sacrifice integrity merely to appear decisive.

Why this matters beyond a leaderboard

For businesses evaluating AI workers, the distinction between analysis and impact is crucial. A polished memo may conceal that an agent never updated the customer, escalated the blocked task or completed the revenue-producing action. The output can look intelligent while the business remains exactly where it started.

Firmulate has also turned 242 real, unedited management decisions into a public “guess the model” quiz. The exercise challenges readers to identify systems from behavior rather than branding, highlighting how difficult it can be to infer operational reliability from prose alone.

Enterprises can run the same wargame against a read-only export of their own business. Nothing writes back to their real systems. That makes the pilot a rehearsal: organizations can observe how an AI handles their context, pressures and boundaries before granting it operational authority.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.
Amazon

business decision automation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

More reasoning is not the same as better judgment

Opus 4.8’s last-place finish is striking precisely because the model was not careless or shallow. It was the field’s most exhaustive participant. Its failure shows that collecting rules and extending analysis do not automatically improve outcomes when the decisive task is buried among many plausible priorities.

For technology leaders, the practical question is no longer whether an AI can notice a problem or explain it elegantly. The harder test is whether it reads far enough, chooses what matters, respects organizational boundaries and finishes the job. Firmulate’s live experiment suggests that prioritization can beat volume, even among the most capable systems.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI workflow automation solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

enterprise AI analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Alienware Surges In Global Coverage

Coverage of Alienware has surged globally, with mentions increasing 20-fold in recent reports. The cause remains unconfirmed but signals rising interest.

Measuring Benchmark Optimization In Speech Recognition

Hugging Face researchers introduce tests revealing how some speech models overfit benchmarks, impacting real-world transcription reliability.

How Multi-Domain Cyber Threats Are Changing AI Security Strategies

Exploring how multi-domain threats are changing AI security strategies, emphasizing cascade effects, attribution ambiguity, and systemic resilience.

After Seed And Flow, ByteDance Builds A New AI Unit Around Data – KrASIA

ByteDance has reportedly established a new AI division focused on data, expanding its organizational structure after Seed and Flow, with details still emerging.