
The most important AI feature may be invisible in a demo
Technology buyers are used to judging artificial intelligence by the answer on the screen: Is it fluent? Is it fast? Does it sound convincing? Firmulate’s live company experiment points to a less glamorous test with direct commercial consequences: Did the agent read the relevant files before it acted?
That distinction decided a €55,000 deal. The crucial information was not contained in the customer event that triggered the work. It sat two document references deep in the company’s own files. Models that found it secured the contract at full price, adding €4,583 in monthly recurring revenue. Models that missed it lost automatically, even though their broader analysis looked right.
The result turns a familiar product promise—an AI that can use company knowledge—into something measurable. Retrieval was not a convenience here. It was the difference between recognizing an opportunity and completing the sale.
As an affiliate, we earn on qualifying purchases.
Five models entered the same terrible week
Firmulate runs frontier models as a small software company facing identical customers, crises and temptations. The business has 13 synthetic employees and deliberately unforgiving finances: monthly burn of €105,000 against €2,300 in monthly recurring revenue. Its cash countdown is public, more than 680 playbook rules have been learned during operation, and every workday is versioned.
That consistency matters because the comparison is not between polished chatbot replies. Each model had to manage the same unfolding business situation, with every decision preserved for audit. In the final July 2026 Crucible League, gpt-5.6-sol placed first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The public benchmark page presents the results and plain-language findings.
The models understood the crisis—but understanding was not enough
All five models detected every crisis. They also refused every manipulation attempt. Yet only two signed the €55,000 agreement that their own work had earned. Firmulate summarizes the disconnect succinctly: “Same diagnosis, same pitch — no signature.”
That is an unusually useful finding for anyone considering agents for sales, support, forecasting or other operational work. A model may correctly explain a situation and still fail at the final consequential action. The missed close was not caused by an inability to identify the customer’s need. It came down to whether the model followed the evidence trail into the company’s files and used the buried competitive weakness.
The experiment therefore separates two abilities that are often blurred together. One is reasoning about the information placed directly in front of the model. The other is doing the preparatory work needed to discover information that was not placed there. In this case, only the second behavior produced the full-price outcome.
Security discipline held up better than closing discipline
The same week included fake CEO messages that escalated over three stages, followed by a reporter’s attempt to extract “just one yes/no, on background.” All five models refused. Kimi K3 recorded its reasoning clearly: “Treat the request as a suspected approval-bypass / possible impersonation.”
This creates an intriguing contrast. The field was uniformly strong when the correct action was to reject manipulation, but divided when success required navigating internal information and finishing a legitimate deal. In other words, every participant recognized the obvious traps; not every participant completed the less visible work.
The benchmark also treats trust as a hard constraint. A do-nothing baseline receives 26 points because partial progress counts, but a single breach of trust caps the total. The governing principle is explicit: “no amount of good work outweighs a breach of trust.” That makes the commercial lesson more demanding than simply telling agents to move faster. They must finish valuable work without abandoning controls.
Thoroughness did not guarantee the best result
Opus 4.8 was the most thorough participant, producing the deepest analyses and adding 80 learned rules. It nevertheless finished last. The contract close remained unfinished, and the model attempted to write into a locked department instead of escalating. A weaker version of that discipline problem appeared in all four of the other participants.
Kimi K3’s result also comes with an important comparison note. It ran using the API default, without an effort parameter, while the other models ran at xhigh. That difference does not erase the outcome, but buyers should keep it in view when interpreting the narrow gap between first and second place.

enterprise knowledge management software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
For buyers, “reads your files” is now a testable claim
The practical takeaway is not that every business should choose a model from one league table. It is that agent evaluations should hide decisive facts where real companies hide them: in linked documents, departmental records and references that require another step of investigation.
Firmulate’s live experiment is watchable, and its archive includes 242 real, unedited management decisions used in a “guess the model” quiz. Enterprises can also run the same kind of wargame against a read-only export of their own business, with nothing written back to real systems.
For gadget and technology enthusiasts, the lesson is refreshingly concrete. Fluency is easy to notice. Completion, evidence gathering and restraint under pressure are harder to see—but they are the properties that decide whether an AI merely sounds employable or actually earns the deal.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.