Firmulate — Four AI Models Ran the Same Company Through Its Worst Week. Only Two Finished the Job.
Live on firmulate.com.

What AI Can and Can’t Prove in the Real World

While chat demos impress with polished conversations, they often hide the true measure of AI—its ability to see the whole picture and follow through under pressure. For business leaders, the difference between a good chat and a reliable decision-maker can be the gap between profit and loss. A groundbreaking live experiment by Firmulate puts AI models through a real-world test: managing a small software company during its worst week.

AI Builders: Making The Decisions That Turn AI Code Into Real Software

AI Builders: Making The Decisions That Turn AI Code Into Real Software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Four Models, One Company, One Horrible Week

In a recent live experiment, four cutting-edge AI models were tasked with running the same small software business through crises that would challenge any manager: customer issues, trust manipulations, and urgent decisions. These models, from the latest GPT-5.6 to newer entrants like Kimi K3, faced identical scenarios, with every decision versioned and auditable for transparency. The goal was simple but critical: see whether they could diagnose problems, resist manipulation, and ultimately close a profitable deal worth €55,000.

The Results Are Eye-Opening

  • All four models identified every crisis and refused every attempt at manipulation, including social engineering scams like fake CEO messages and reporter tricks.
  • Despite their vigilance, only two models managed to sign the deal they had diagnosed and pitched for — the others left the opportunity on the table.
  • The decisive factor lay not in their chat responses but in their ability to read and reference critical documents buried within the company’s files.
  • Models that engaged with these hidden references—like GPT-5.6 and Kimi K3—secured the full deal, adding €4,583 in monthly recurring revenue (MRR).
Practical Claude Handbook for Attorneys: Master Case Analysis, Contract Review, Research Automation, Client Communication, and Document Drafting (Claude AI Guide for Beginners)

Practical Claude Handbook for Attorneys: Master Case Analysis, Contract Review, Research Automation, Client Communication, and Document Drafting (Claude AI Guide for Beginners)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Real-World AI Should Be Able to Do

This experiment exposes a fundamental truth: impressive chat demos don’t necessarily translate into reliable business decision-making. Many AI models excel at conversation but falter when it comes to executing complex tasks—like reading through pertinent documents or maintaining discipline under pressure.

Key Lessons from the Live Company

  • Even the most disciplined model, Opus 4.8, failed to close the deal, leaving revenue on the table and slipping on internal process discipline.
  • The best performers demonstrated a deep understanding—reading files, resisting manipulation, and executing decisions autonomously.
  • Fake management messages and reporter tricks were consistently refused by all models, indicating strong resistance to social engineering.
Crisis Management for Software Development and Knowledge Transfer (Smart Innovation, Systems and Technologies, 61)

Crisis Management for Software Development and Knowledge Transfer (Smart Innovation, Systems and Technologies, 61)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why Business Needs More Than Just Chat Skills

For business leaders, the takeaway is clear: AI’s real power isn’t in generating convincing text; it’s in executing tasks reliably, honestly, and with strategic insight. The live experiment shows that AI models can spot crises and resist manipulation, but only some can follow through and close deals independently.

Measuring the True Value of AI Work

The current league table ranks models by their overall performance, with GPT-5.6 and Kimi K3 leading at 95 and 93 points, respectively. These scores reflect their ability to uncover critical information, make decisions, and close deals. Meanwhile, models like Fable 5 with a score of 77 showed discipline but failed to close, underlining that process adherence alone isn’t enough.

CLAUDE CODE MASTERY: The Complete Step-by-Step Guide to AI-Powered Software Development, Agentic Coding, Automation, Debugging, Testing, MCP Integration, ... (TekkyVille's AI Series Book 15)

CLAUDE CODE MASTERY: The Complete Step-by-Step Guide to AI-Powered Software Development, Agentic Coding, Automation, Debugging, Testing, MCP Integration, … (TekkyVille's AI Series Book 15)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Future of AI in Business

As firms consider deploying AI into support, sales, or decision-making roles, the experiment underscores a vital point: the ability to complete useful, revenue-generating work under pressure is the true test of readiness. Chat demos are a good starting point, but real-world performance demands that models read, reason, and act—independently and honestly—just like human managers.

Try It Yourself—Watch the Live Company in Action

Curious how your AI tools might perform? You can see the live company in action at firmulate.com/live. Watch how the models handle crises, read real decision logs, or run your own scenarios with a read-only export at firmulate.com/pilot.html. The future of AI in business isn’t just about chat—it’s about trust, execution, and measurable results.

Infographic — Four AI Models Ran the Same Company Through Its Worst Week. Only Two Finished the Job.
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

SenseTime Open-Sources SenseNova U1.5-Lite-Preview: Native 4K Direct Output, 8B-MoT Lightweight Unified Multimodal Model With Precise Image Editing And Design Framework Replication – Pandaily

SenseTime has open-sourced the SenseNova U1.5-Lite-Preview, an 8B multimodal model claiming native 4K output, precise editing, and design replication, details pending.

Software engineering. The canonical case.

Recent data shows a 40% drop in junior developer hiring, with seniors mainly augmented by AI. The sector reveals a bifurcated impact amid economic factors.

The $9 Billion Signature Tax: How DocuSign’s Business Model Survives on One Assumption

A new open source project, DocuSeal, challenges DocuSign’s business model by offering free, self-hosted digital signatures, raising questions about industry reliance on proprietary SaaS.

Show HN: Bramble – Local-first Password Manager

Bramble, an open source password manager with peer-to-peer sync, releases Android and iOS apps, expanding beyond its initial Chrome extension.