AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

A live experiment by Firmulate tested AI models managing a small company during its worst week. The results reveal that management skills, not just chat quality, should be the new benchmark for AI evaluation. The top model scored 95, but all models faced challenges in trust and execution.

In a groundbreaking live experiment, Firmulate has tested five AI models managing a small software company’s worst week, revealing that management quality surpasses chat responsiveness as the key metric for AI evaluation.

The experiment, called the Crucible League 2026, ranked models based on their ability to diagnose crises, communicate effectively, and maintain trust under pressure. For more details, see the original analysis. GPT-5.6-SOL led with a score of 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77, and Opus 4.8 with 73. The models were tasked with managing a simulated crisis environment involving real money mechanics, customer interactions, and operational decisions. This kind of testing is detailed in the original analysis.

While all models identified crises and resisted manipulation attempts, only two successfully closed a €55,000 deal based on their analysis. This highlights the importance of effective management, as discussed in the internal coverage. The experiment emphasized that effective management involves more than generating plausible responses; it requires retrieving critical facts, making decisions, and executing them reliably. Notably, models that produced more detailed analyses did not necessarily perform better in execution, highlighting a gap between effort and effective management.

At a glance
reportWhen: ongoing, with results finalized in July…
The developmentFirmulate’s live management experiment evaluates AI models’ ability to handle real-world business crises, exposing gaps in current AI benchmarks.

Why Management Skills Outperform Chat Quality in AI Evaluation

This experiment demonstrates that current AI benchmarks, which often focus on language fluency or technical output, are insufficient for real-world management tasks. The ability to triage crises, maintain trust, and complete decisions is crucial for deploying AI in operational roles. The results suggest that AI evaluation frameworks should incorporate management-oriented metrics, emphasizing trustworthiness, decision accuracy, and execution reliability. For organizations considering AI assistants, this shift could mean prioritizing models capable of managing organizational consequences rather than just generating convincing responses.

Amazon

AI management simulation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Revealing the Limitations of Traditional AI Benchmarks

Existing AI benchmarks primarily assess models on coding, language understanding, or conversational quality. These tests often overlook the complexities of managing real-world scenarios, especially under pressure. The Firmulate experiment builds on the growing recognition that effective AI in business must handle crises, prioritize tasks, and preserve trust over time. Previous efforts have tested AI in isolated tasks, but this live management scenario offers a more comprehensive and realistic evaluation environment. The July 2026 Crucible League final results provide a rare glimpse into how models perform when responsible for actual organizational outcomes, not just responses.

“Management quality, not chat quality, should define AI benchmarks. Our live test shows models can identify crises but struggle with execution and trust.”

— Thorsten Meyer, creator of the experiment

Amazon

business crisis management AI software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unresolved Questions About AI Management Performance

While the experiment clearly ranks models based on their management capabilities, it remains unclear how these results translate to diverse real-world organizations. The long-term reliability of these models under different operational contexts, and their ability to handle unforeseen crises, are still being studied. Additionally, the impact of training data, model architecture, and deployment environment on management effectiveness needs further investigation. It is also uncertain whether current evaluation metrics fully capture the nuanced qualities of trust, ethical judgment, and strategic decision-making in complex business settings.

Amazon

AI decision-making tools for organizations

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for AI Management Benchmarks and Adoption

Organizations interested in deploying AI for management roles should consider conducting internal simulations or wargames similar to the Firmulate experiment to evaluate model performance in their specific context. Researchers and developers are expected to refine evaluation metrics, integrating management-oriented tasks that assess decision quality, trustworthiness, and execution reliability. Meanwhile, the AI community may see new benchmarks emerge that focus on managing consequences rather than just generating high-quality responses. The ongoing analysis of these models’ real-world applications will shape industry standards and guide responsible adoption of AI in operational roles.

Amazon

AI trustworthiness evaluation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Why is management quality more important than chat performance for AI models?

Management quality reflects an AI’s ability to handle crises, make decisions, and maintain trust, which are critical for operational success. Chat performance alone does not measure these capabilities.

How does the Firmulate experiment test AI models’ management skills?

It simulates managing a company’s worst week, requiring models to diagnose crises, communicate decisions, and execute tasks under pressure, with real financial and operational stakes involved.

What are the main limitations of current AI benchmarks based on this experiment?

They often focus on language fluency or technical correctness, neglecting the ability to manage organizational consequences, build trust, and execute decisions reliably in real-world scenarios.

Will these findings influence how companies evaluate AI models in the future?

Yes, the results suggest that management-oriented metrics will become increasingly important, prompting organizations to develop or adopt new evaluation frameworks that emphasize decision-making and trustworthiness.

What should organizations do before deploying AI for management tasks?

They should conduct internal simulations or wargames to test AI models in their specific operational environment, assessing not just response quality but also decision accuracy, trust, and execution reliability.

Source: ThorstenMeyerAI.com

FALL YARD WORK

Fall yard work Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

VigilSAR: The Object That Isn’t Transmitting

VigilSAR uses SAR and data fusion to identify ships that operate without transponders, enhancing maritime awareness in all weather conditions.

AI Black Boxes And The Fragile Future Of Global Alliances

Emerging AI black boxes threaten NATO’s supply chains and alliances, raising concerns over control, security, and dependency in critical infrastructure.

NTT DATA Group Cuts Incident Analysis To 30 Minutes With Codex

NTT DATA Group reports reducing incident analysis time to 30 minutes with OpenAI Codex, but details on scope, baseline, and impact remain unclear.

Xbox Outage

An ongoing Xbox outage has affected players nationwide, with Microsoft confirming service disruptions. Details on cause and resolution are still emerging.