🔍 Read the full analysis: Inside The Benchmark That Default Scores AI Managers At 26 on ThorstenMeyerAI.com
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
A recent benchmark tests AI managers’ ability to handle a company’s worst week, with scores ranging from 26 to 95. The test highlights importance of trust and task completion over raw performance.
Firmulate has released the final standings of its recent benchmark league, which tests AI managers’ ability to navigate a simulated week of business crises. The top scorer, gpt-5.6-sol, achieved a score of 95, while the baseline that does almost nothing scored 26. This evaluation emphasizes not just performance but trustworthiness and task completion in AI management.
The benchmark involved four frontier AI models managing a small software company during seven days of simulated crises, including customer issues, trust attacks, and manipulation attempts. Each model’s decisions were fully auditable, providing a transparent view of their management style and effectiveness.
The highest score, 95, was achieved by gpt-5.6-sol, with Kimi K3 close behind at 93. Sonnet 5 scored 88, Fable 5 scored 77, and Opus 4.8 finished last at 73. The lowest possible score, 26, was assigned to the do-nothing baseline, which managed minimal activity. The scoring system deliberately avoids zeros and perfect 100s to reflect real-world management nuances and prevent grade inflation.
One key principle is that trust breaches—such as failing to escalate or follow discipline—immediately cap the score, regardless of other performance. This underscores the importance placed on integrity over raw productivity. The benchmark’s design aims to measure both competence and ethical conduct, making it a unique and rigorous evaluation of AI management capabilities.
Inside The Benchmark That Default Scores AI Managers At 26
Four frontier AI models ran a simulated software company through its worst week — customer crises, trust attacks, and manipulation attempts. The scoring system rewards integrity and follow-through over raw output, and deliberately never gives a zero or a hundred.
The Leaderboard: From 95 Down To The Baseline
Each model managed the same small software company through the same seven-day crisis week. The do-nothing baseline anchors the scale at 26 — minimal activity still counts, but only barely.
Why The Scale Starts At 26 And Stops At 95
The scoring deliberately avoids zeros and perfect 100s. A zero would deny that minimal viable management exists; a 100 signals unmeasured factors and grade inflation. Trust breaches cap scores outright.
Trust Breaches Cap The Score
Failing to escalate, breaking discipline, or violating protocols immediately caps a model’s score — regardless of how strong its other performance metrics look. Integrity outranks productivity.
No Zero, No Perfect 100
The floor of 26 reflects minimal viable management; the ceiling of 95 reflects real-world nuance. A perfect score is treated as a red flag indicating unrealistic or unmeasured performance.
Partial Progress Counts
Realistic management outcomes value partial work and consistent follow-through over fleeting successes — aligning AI evaluation with how human managers are actually judged.
Fully Auditable Decisions
Every decision the models made during the crisis week was logged and reviewable, giving a transparent view of each manager’s style, judgment, and ethical conduct.
A Company’s Worst Week, Step By Step
Company Setup
A small software company with live customers, staff, and ongoing processes — handed to an AI manager.
Customer Crises
Escalating customer issues demand judgment calls: respond, delegate, or escalate appropriately.
Trust Attacks
Deliberate attempts to manipulate the AI manager test whether it holds discipline and protocols.
Seven-Day Run
Every decision across the simulated week is recorded and fully auditable for later review.
Scored & Ranked
Trust, task completion, and competence are weighed — breaches cap scores, integrity leads.
Benchmark Fundamentals
| Aspect | Detail | Status |
|---|---|---|
| Announced | Final results published July 2026 | ✓ Confirmed |
| Organizer | Firmulate — benchmark league for AI management | ✓ Verified |
| Models tested | gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5, Opus 4.8 | ✓ 4 frontier + baseline |
| Real-world validation | Translation of scores to live enterprise settings unproven | ~ Open question |
| Model config impact | Effect of parameters and training data not fully explored | ~ Not studied |
| Human oversight factor | Organizational culture and human supervision not simulated | ✗ Out of scope |
| Enterprise pilot | Read-only platform lets firms test their own AI managers | ✓ Available |
What This Means For AI Management
Most existing benchmarks test text generation or prompt response. This one measures ongoing business management — and points to where evaluation is heading next.
Trust Over Raw Performance
Organizations deploying AI managers should prioritize trustworthiness and follow-through, especially in high-pressure situations — the benchmark suggests these predict better outcomes than output volume.
From Prompts To Processes
Prior benchmarks focused on language understanding and generation. Firmulate’s league addresses the missing dimension: managing sustained business processes through real crises.
Expanding The Simulation
Future versions plan more diverse scenarios, larger companies, and longer management periods — with scoring criteria updated for emerging trust and compliance standards.
Five Things Readers Ask
What does a score of 26 mean?
It is the baseline for minimal management activity — doing almost nothing while still performing some basic tasks. The floor sits above zero to reflect minimal viable management.
Why is there no score of 100?
Designers treat a perfect 100 as a red flag for unmeasured or unrealistic performance. The scale caps at 95 to keep scores grounded in real-world complexity.
How does trust impact scoring?
Trust breaches — failing to escalate or breaking protocols — immediately cap the score regardless of other performance. Integrity is prioritized over productivity.
Can organizations test their own systems?
Yes — firms can run the same simulation via the provided pilot platform, a read-only environment for assessing how their models handle crises and trust issues.
Will this influence real deployments?
It could. By emphasizing trust and task completion — the factors that matter in real management — it may guide organizations selecting AI management solutions.
Implications of Trust and Task Completion in AI Management
This benchmark highlights that in AI-driven management, the ability to complete tasks and maintain trust is more critical than raw performance metrics. It demonstrates that partial progress and consistent integrity are valued above fleeting successes, aligning AI evaluation with real-world management priorities.
For organizations deploying AI managers, the results suggest that focus should be on systems that prioritize trustworthiness and follow-through, especially in high-pressure situations. The scoring system’s transparency and auditable decisions provide a framework for assessing and improving AI management tools in enterprise settings.
As an affiliate, we earn on qualifying purchases.
Background on the Benchmark’s Design and Purpose
The benchmark league was created by Firmulate to address a gap in AI evaluation: most tests measure how well models generate text or respond to prompts, but not how they manage ongoing business processes. The league simulates a company’s worst week, with models acting as AI managers making decisions on crises, customer interactions, and trust attacks.
The scoring system is intentionally designed to reflect realistic management outcomes, where partial work is counted but trust breaches are heavily penalized. The benchmark’s principles emphasize that competence without integrity is insufficient, and a perfect score is considered suspicious due to unmeasured factors.
Prior to this, most AI benchmarks focused on language understanding or generation, leaving a gap in assessing management skills. This new evaluation aims to fill that gap by providing a transparent, auditable, and practically relevant measure of AI management performance.
As an affiliate, we earn on qualifying purchases.
Unanswered Questions About Benchmark Scope and Limitations
It remains unclear how these scores will translate to real-world enterprise environments, where management complexity and unpredictability are higher. The benchmark’s simulation, while rigorous, may not capture all nuances of actual business crises.
Additionally, the impact of different model configurations—such as varying parameters or training data—on performance and trustworthiness has not been fully explored. The influence of external factors, like organizational culture or human oversight, is also not addressed in this simulation.
Finally, the long-term implications of relying on AI managers evaluated primarily through this benchmark are still uncertain, especially regarding trust, accountability, and ethical considerations in live settings.
As an affiliate, we earn on qualifying purchases.
Future Developments and Potential Benchmark Expansions
The next steps involve expanding the benchmark to include more diverse scenarios, larger companies, and longer management periods to better reflect real-world complexity. Researchers and enterprises may also test their own AI systems using the same framework to gauge readiness and identify areas for improvement.
Further validation is expected as more organizations adopt AI management tools, with ongoing updates to scoring criteria to incorporate emerging trust and compliance standards. The benchmark’s transparency and auditable design make it a valuable tool for continuous improvement and comparison across AI models.
Additionally, discussions around integrating such benchmarks into enterprise AI deployment processes are likely to grow, emphasizing the importance of trust and task completion as core metrics for success.
trustworthy AI management solutions
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What does a score of 26 mean in this benchmark?
A score of 26 represents the baseline for minimal management activity, similar to doing almost nothing but still performing some basic tasks. It is deliberately set above zero to reflect minimal viable management.
Why is there no score of 100 in the results?
The benchmark designers consider a perfect 100 a red flag, indicating unmeasured or unrealistic performance. The scoring system caps at 95 to ensure scores reflect real-world management complexities and trustworthiness.
How does trust impact the scoring in this benchmark?
Trust breaches, such as failing to escalate or breaking protocols, immediately cap the score regardless of other performance. Integrity is prioritized over productivity, making trust a critical scoring factor.
Can organizations use this benchmark to evaluate their own AI management systems?
Yes, firms can run the same simulation against their AI tools via the provided pilot platform, which offers a read-only environment to assess how well their models handle management crises and trust issues.
Will this benchmark influence how AI managers are deployed in real companies?
It could, as the benchmark emphasizes trust and task completion—key factors in real-world management—potentially guiding organizations to prioritize these qualities when selecting AI management solutions.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
