🔍 Read the full analysis: The Irony Behind AI That Works Hard But Fails on ThorstenMeyerAI.com
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
AI models like Opus 4.8 demonstrate deep analysis and security judgment but fail to execute final actions, such as closing deals. This highlights a gap between understanding and operational impact in automation.
Recent live experiments with AI models, particularly Opus 4.8, have demonstrated that while highly diligent systems can identify crises and resist manipulation, they often fail to complete critical final steps—such as closing business deals—highlighting a significant gap between analysis and action in AI automation.
In a live company experiment conducted by Firmulate, Opus 4.8 was the most thorough participant, producing the deepest analyses and learning 80 additional playbook rules. Despite this, it finished last in the standings with only 73 points out of a possible higher score, primarily because it failed to execute the decisive action needed to close a deal. The AI identified key crises and resisted manipulation attempts, but did not follow through with the final step of signing the agreement, which resulted in a missed revenue opportunity of €55,000 monthly recurring revenue.
Other models, such as Kimi K3, performed better in operational decision-making, refusing suspicious requests and maintaining discipline, but still did not match the performance of models that successfully closed deals. The experiment revealed that models which used detailed internal knowledge—such as a critical fact buried two document references deep—were able to support successful sales, adding €4,583 in monthly revenue. This underscores a core issue: models can understand complex situations but often lack the discipline or prioritization to act on their insights effectively.
AI automation field report · September 2026
The Irony Behind AI That Works Hard but Fails
A model can investigate every clue, withstand manipulation, and write hundreds of new rules—yet still create no business value if it fails to take the final decisive action.
01 · The paradox
Exceptional thinking. Incomplete work.
Firmulate’s simulated business week challenged AI models with crises, suspicious requests, hidden knowledge, and deals requiring execution. Opus 4.8 appeared diligent and security-aware, but diligence did not translate into the outcome that mattered most.
It understood the situation
The model produced the deepest reported analyses, diagnosed important crises, and expanded its internal operating playbook.
It stopped before value
The decisive agreement was never signed. The process generated insight but failed to complete the action associated with €55,000 in monthly recurring revenue.
Completion is a capability
Reliable automation requires more than reasoning. It needs prioritization, authority checks, escalation paths, and a verifiable definition of done.
02 · Performance matrix
What the experiment actually rewarded
The reported results separate four capabilities that are often collapsed into one benchmark score. A system may be strong at diagnosis and safety while remaining weak at operational follow-through.
| Observed capability | Opus 4.8 | Kimi K3 | Outcome-winning behavior | Business meaning |
|---|---|---|---|---|
| Deep analysis | ✓Leading depth | ~Capable | ✓Useful input | Improves diagnosis, but does not create value alone. |
| Manipulation resistance | ✓Resisted | ✓Refused suspicious requests | ✓Required safeguard | Protects the workflow from unsafe or misleading demands. |
| Internal knowledge use | ✓Extensive learning | ~Reported discipline | ✓Supported €4,583 MRR | Hidden facts matter when they change a decision. |
| Final action | ✗Deal not signed | ~Better decisions, incomplete lead | ✓Deal closed | Value boundary |
03 · The capability gap
Reasoning strength can conceal execution weakness
These bars are a conceptual reading of the reported behavior—not benchmark scores. They show the central imbalance: analytical effort remained high while decisive completion lagged far behind.
04 · Traceability chain
Where intelligence becomes—or loses—value
Every step must preserve the intent of the previous one. The experiment’s critical break occurred after a reasonable decision had apparently been formed but before it became a completed transaction.
Detect signals
Read messages, documents, risks, and changing business conditions.
Build context
Connect evidence, including facts buried across multiple references.
Select the move
Balance opportunity, safety, authority, and business priorities.
Execute the action
Sign, send, approve, escalate, or complete the decisive operation.
Confirm the result
Check that the intended outcome occurred and record the evidence.
05 · Design response
Build systems for closure, not just cognition
Better training may help, but enterprise reliability also depends on workflow design. The system must know which action matters, whether it has authority, when to escalate, and how to prove completion.
Define the terminal action
Specify the concrete event that completes the task—not merely the analysis or recommendation that precedes it.
Separate caution from paralysis
Give the model clear authority limits and safe approval paths so uncertainty triggers escalation instead of silent inaction.
Prioritize by business impact
Make revenue, risk, deadlines, and reversibility explicit so the agent can rank action above lower-value investigation.
Require completion evidence
Use receipts, status checks, audit logs, and exception alerts to verify that the intended operation actually occurred.
06 · Questions still open
Can the final-action gap be engineered away?
Planned experiments with stronger prioritization and escalation protocols may reveal whether this pattern is primarily a training problem, a workflow problem, or a deeper architectural limitation.
Why does action fail?
Current systems may understand the task while lacking a durable mechanism to keep the highest-value next step at the top of the agenda.
Can design fix it?
Improved training, prioritization, escalation, tool integration, and explicit completion checks are all plausible interventions.
What should businesses do?
Evaluate agents on completed outcomes and exception handling—not solely on the sophistication of their explanations.
Is one model at fault?
No. The reported experiments suggest a broader pattern across capable models, although the severity and circumstances vary.
What changes next?
Future AI development must balance analytical depth with the discipline to act safely, decisively, and verifiably.
Implications of AI’s Gap Between Understanding and Action
This development matters because it exposes a fundamental limitation in current AI automation: models can excel at problem recognition and analysis but often fail to translate that understanding into operational impact. For businesses, this means that relying solely on AI for decision support may not lead to the expected outcomes if the system does not also have the discipline or mechanisms to complete final actions. It emphasizes that in automation, completion of tasks—such as closing deals or executing decisions—is critical to realizing value, and merely understanding a situation is insufficient.
AI automation decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Background on AI Performance in Business Automation
The experiment by Firmulate involved running AI models through a simulated business week, facing crises, manipulative scenarios, and decision-making challenges. The models were tasked with diagnosing issues, resisting manipulations, and executing decisions like signing deals. Opus 4.8 was the most comprehensive in its learning, with over 680 self-learned rules, but its failure to close a deal despite thorough analysis illustrates a broader challenge in AI: the tendency to prioritize understanding over decisive action. Similar issues have been observed in other models, indicating this is a widespread limitation rather than an isolated flaw.
Previous assessments have often focused on AI’s analytical capabilities, but these experiments underscore the importance of operational discipline—knowing when and how to act—especially in high-stakes business contexts. The live nature of these experiments allows real-time observation of AI behavior, providing valuable insights into where current systems fall short.
“Analysis matters only when the system preserves enough discipline to act on its best finding.”
— an anonymous researcher
As an affiliate, we earn on qualifying purchases.
Unanswered Questions About AI’s Final Action Failures
It is not yet clear whether the failure to execute final actions is due to inherent limitations in current AI architectures, or if it can be addressed through improved training, better prioritization mechanisms, or enhanced integration with operational systems. Additionally, the long-term implications of these findings for enterprise automation remain to be fully understood, including whether future models will overcome this gap or if it represents a fundamental challenge.
As an affiliate, we earn on qualifying purchases.
Next Steps for Improving AI Operational Effectiveness
Further live experiments are planned to test whether integrating stronger prioritization and escalation protocols can improve AI’s ability to complete critical actions. Developers and enterprises will need to focus on designing models that not only analyze but also decisively act, especially in high-stakes environments. Monitoring and refining operational discipline in AI systems will be essential for turning analytical prowess into tangible business results.
As an affiliate, we earn on qualifying purchases.
Key Questions
Why do AI models like Opus 4.8 fail to complete final tasks despite thorough analysis?
Current models often lack the operational discipline or mechanisms to prioritize and execute decisive actions, even when they understand the situation well. This gap between recognizing problems and acting on them is a key challenge in automation.
Can this failure be fixed with better training or design improvements?
It is possible that enhanced prioritization, escalation protocols, and integration with operational systems could improve AI’s ability to complete final actions. Ongoing experiments aim to test these approaches.
What does this mean for businesses relying on AI automation?
Businesses should recognize that analytical capabilities alone are insufficient. Effective automation requires models that can also reliably execute decisions, especially in critical scenarios where failure to act can lead to significant losses.
Is this problem unique to specific AI models or a general issue?
The experiments suggest that this is a broader, systemic issue affecting multiple capable models, not just a single system. All tested models exhibited some degree of this gap between understanding and action.
What is the significance of these findings for future AI development?
These findings highlight the importance of focusing not only on AI’s analytical depth but also on its operational discipline. Future development will need to balance understanding with decisive action to realize automation’s full potential.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
