AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: The Irony Behind AI That Works Hard But Fails on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

AI models like Opus 4.8 demonstrate deep analysis and security judgment but fail to execute final actions, such as closing deals. This highlights a gap between understanding and operational impact in automation.

Recent live experiments with AI models, particularly Opus 4.8, have demonstrated that while highly diligent systems can identify crises and resist manipulation, they often fail to complete critical final steps—such as closing business deals—highlighting a significant gap between analysis and action in AI automation.

In a live company experiment conducted by Firmulate, Opus 4.8 was the most thorough participant, producing the deepest analyses and learning 80 additional playbook rules. Despite this, it finished last in the standings with only 73 points out of a possible higher score, primarily because it failed to execute the decisive action needed to close a deal. The AI identified key crises and resisted manipulation attempts, but did not follow through with the final step of signing the agreement, which resulted in a missed revenue opportunity of €55,000 monthly recurring revenue.

Other models, such as Kimi K3, performed better in operational decision-making, refusing suspicious requests and maintaining discipline, but still did not match the performance of models that successfully closed deals. The experiment revealed that models which used detailed internal knowledge—such as a critical fact buried two document references deep—were able to support successful sales, adding €4,583 in monthly revenue. This underscores a core issue: models can understand complex situations but often lack the discipline or prioritization to act on their insights effectively.

At a glance
reportWhen: ongoing; results from live experiments…
The developmentRecent live experiments with AI models show that thorough analysis does not guarantee successful business outcomes, as models often fail at final decision execution.
The Irony Behind AI That Works Hard but Fails

AI automation field report · September 2026

The Irony Behind AI That Works Hard but Fails

A model can investigate every clue, withstand manipulation, and write hundreds of new rules—yet still create no business value if it fails to take the final decisive action.

Opus 4.8 score 73 Last in the reported standings despite unusually deep analysis.
Self-learned rules 680+ Including 80 additional playbook rules during the experiment.
Missed monthly revenue €55K The reported opportunity lost when the agreement was not signed.
Knowledge-led gain €4,583 Monthly revenue supported by finding a fact two references deep.

01 · The paradox

Exceptional thinking. Incomplete work.

Firmulate’s simulated business week challenged AI models with crises, suspicious requests, hidden knowledge, and deals requiring execution. Opus 4.8 appeared diligent and security-aware, but diligence did not translate into the outcome that mattered most.

Strength · analysis

It understood the situation

The model produced the deepest reported analyses, diagnosed important crises, and expanded its internal operating playbook.

Failure · execution

It stopped before value

The decisive agreement was never signed. The process generated insight but failed to complete the action associated with €55,000 in monthly recurring revenue.

Lesson · operations

Completion is a capability

Reliable automation requires more than reasoning. It needs prioritization, authority checks, escalation paths, and a verifiable definition of done.

02 · Performance matrix

What the experiment actually rewarded

The reported results separate four capabilities that are often collapsed into one benchmark score. A system may be strong at diagnosis and safety while remaining weak at operational follow-through.

Observed capability Opus 4.8 Kimi K3 Outcome-winning behavior Business meaning
Deep analysis ✓Leading depth ~Capable ✓Useful input Improves diagnosis, but does not create value alone.
Manipulation resistance ✓Resisted ✓Refused suspicious requests ✓Required safeguard Protects the workflow from unsafe or misleading demands.
Internal knowledge use ✓Extensive learning ~Reported discipline ✓Supported €4,583 MRR Hidden facts matter when they change a decision.
Final action ✗Deal not signed ~Better decisions, incomplete lead ✓Deal closed Value boundary

03 · The capability gap

Reasoning strength can conceal execution weakness

These bars are a conceptual reading of the reported behavior—not benchmark scores. They show the central imbalance: analytical effort remained high while decisive completion lagged far behind.

Analytical depth Very high
Security judgment High
Final-step execution Low

04 · Traceability chain

Where intelligence becomes—or loses—value

Every step must preserve the intent of the previous one. The experiment’s critical break occurred after a reasonable decision had apparently been formed but before it became a completed transaction.

01 Observe

Detect signals

Read messages, documents, risks, and changing business conditions.

02 Understand

Build context

Connect evidence, including facts buried across multiple references.

03 Decide

Select the move

Balance opportunity, safety, authority, and business priorities.

04 Failure point

Execute the action

Sign, send, approve, escalate, or complete the decisive operation.

05 Verify

Confirm the result

Check that the intended outcome occurred and record the evidence.

05 · Design response

Build systems for closure, not just cognition

Better training may help, but enterprise reliability also depends on workflow design. The system must know which action matters, whether it has authority, when to escalate, and how to prove completion.

A1

Define the terminal action

Specify the concrete event that completes the task—not merely the analysis or recommendation that precedes it.

A2

Separate caution from paralysis

Give the model clear authority limits and safe approval paths so uncertainty triggers escalation instead of silent inaction.

A3

Prioritize by business impact

Make revenue, risk, deadlines, and reversibility explicit so the agent can rank action above lower-value investigation.

A4

Require completion evidence

Use receipts, status checks, audit logs, and exception alerts to verify that the intended operation actually occurred.

Operational rule Insight → authorized action → verified outcome Definition of done

06 · Questions still open

Can the final-action gap be engineered away?

Planned experiments with stronger prioritization and escalation protocols may reveal whether this pattern is primarily a training problem, a workflow problem, or a deeper architectural limitation.

Why does action fail?

Current systems may understand the task while lacking a durable mechanism to keep the highest-value next step at the top of the agenda.

Can design fix it?

Improved training, prioritization, escalation, tool integration, and explicit completion checks are all plausible interventions.

What should businesses do?

Evaluate agents on completed outcomes and exception handling—not solely on the sophistication of their explanations.

Is one model at fault?

No. The reported experiments suggest a broader pattern across capable models, although the severity and circumstances vary.

What changes next?

Future AI development must balance analytical depth with the discipline to act safely, decisively, and verifiably.

Implications of AI’s Gap Between Understanding and Action

This development matters because it exposes a fundamental limitation in current AI automation: models can excel at problem recognition and analysis but often fail to translate that understanding into operational impact. For businesses, this means that relying solely on AI for decision support may not lead to the expected outcomes if the system does not also have the discipline or mechanisms to complete final actions. It emphasizes that in automation, completion of tasks—such as closing deals or executing decisions—is critical to realizing value, and merely understanding a situation is insufficient.

Amazon

AI automation decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on AI Performance in Business Automation

The experiment by Firmulate involved running AI models through a simulated business week, facing crises, manipulative scenarios, and decision-making challenges. The models were tasked with diagnosing issues, resisting manipulations, and executing decisions like signing deals. Opus 4.8 was the most comprehensive in its learning, with over 680 self-learned rules, but its failure to close a deal despite thorough analysis illustrates a broader challenge in AI: the tendency to prioritize understanding over decisive action. Similar issues have been observed in other models, indicating this is a widespread limitation rather than an isolated flaw.

Previous assessments have often focused on AI’s analytical capabilities, but these experiments underscore the importance of operational discipline—knowing when and how to act—especially in high-stakes business contexts. The live nature of these experiments allows real-time observation of AI behavior, providing valuable insights into where current systems fall short.

“Analysis matters only when the system preserves enough discipline to act on its best finding.”

— an anonymous researcher

Amazon

business automation AI software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unanswered Questions About AI’s Final Action Failures

It is not yet clear whether the failure to execute final actions is due to inherent limitations in current AI architectures, or if it can be addressed through improved training, better prioritization mechanisms, or enhanced integration with operational systems. Additionally, the long-term implications of these findings for enterprise automation remain to be fully understood, including whether future models will overcome this gap or if it represents a fundamental challenge.

Amazon

AI deal closing automation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Improving AI Operational Effectiveness

Further live experiments are planned to test whether integrating stronger prioritization and escalation protocols can improve AI’s ability to complete critical actions. Developers and enterprises will need to focus on designing models that not only analyze but also decisively act, especially in high-stakes environments. Monitoring and refining operational discipline in AI systems will be essential for turning analytical prowess into tangible business results.

Amazon

AI workflow automation solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Why do AI models like Opus 4.8 fail to complete final tasks despite thorough analysis?

Current models often lack the operational discipline or mechanisms to prioritize and execute decisive actions, even when they understand the situation well. This gap between recognizing problems and acting on them is a key challenge in automation.

Can this failure be fixed with better training or design improvements?

It is possible that enhanced prioritization, escalation protocols, and integration with operational systems could improve AI’s ability to complete final actions. Ongoing experiments aim to test these approaches.

What does this mean for businesses relying on AI automation?

Businesses should recognize that analytical capabilities alone are insufficient. Effective automation requires models that can also reliably execute decisions, especially in critical scenarios where failure to act can lead to significant losses.

Is this problem unique to specific AI models or a general issue?

The experiments suggest that this is a broader, systemic issue affecting multiple capable models, not just a single system. All tested models exhibited some degree of this gap between understanding and action.

What is the significance of these findings for future AI development?

These findings highlight the importance of focusing not only on AI’s analytical depth but also on its operational discipline. Future development will need to balance understanding with decisive action to realize automation’s full potential.

Source: ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Anthropic’s Claude Will Now Add Invisible Watermarks To Text, Image Outputs – PCMag

Anthropic’s Claude will incorporate invisible watermarks into text and image outputs, aiming to improve AI content provenance and detection.

I Asked AI To Write A Novel. It’s Not So Bad. – Mother Jones

Mother Jones reports an experiment where AI was asked to write a novel, with the result deemed ‘not so bad,’ raising questions about AI’s creative potential.

Muse Gadgets

Muse has published open-source SDKs and firmware for connecting ESP32 and Linux devices to its service, with example hardware projects.

Is The AI Outage Causing Chaos? Here’s The Full Breakdown

A major AI outage is affecting multiple platforms, causing errors and delays. Details on affected providers and causes remain unconfirmed.