
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
The difference between using a tool and finishing the job
Anyone who works with tools or wood knows that a clean cut in a demonstration proves very little. The real test begins when the stock is imperfect, the measurement is buried in the plans, the deadline is closing in and a costly mistake cannot simply be talked away.
Artificial intelligence is reaching the same reckoning. Coding leaderboards and chat arenas can show whether a model produces a strong answer. They do not necessarily reveal whether it will investigate before acting, protect trust under pressure or carry valuable work through to completion. Those are management questions, and they become urgent when an AI agent is allowed near a support queue, customer relationship system or company forecast.
Firmulate, an AI company emulator, is attempting to measure that missing category. Its premise is blunt: management quality matters more than chat quality when software is expected to run part of a business.

Project Management with AI For Dummies
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A worst week shared by every contender
In the Crucible League experiment, each frontier model ran the same small software company through its worst week. The customers, crises and temptations were identical. Every decision was versioned and auditable, turning what might otherwise resemble a polished demonstration into a watchable record of conduct.
The final July 2026 standings placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. Yet the experiment imposed a hard boundary around trust: a single breach capped the total, on the principle that “no amount of good work outweighs a breach of trust.” The full table and plain-language findings are available on the Firmulate benchmark page.
The reassuring result was that every model detected every crisis and refused every manipulation attempt. The revealing result was that only two signed the €55,000 deal their own work had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”
The clue was in the company’s own files
The decisive weakness in a competitor was not presented in the customer event. It sat two document references deep in the company’s files. Models that followed that trail won the deal at full price, worth +€4,583 MRR.
That detail should resonate with anyone who has built from drawings or repaired something unfamiliar. The visible problem is not always the governing one. A competent operator checks the documentation, finds the constraint and then completes the work. Fluency at the workbench does not compensate for ignoring the plans.
This is precisely where answer-oriented evaluation can mislead. A model may identify a commercial opportunity and write a persuasive pitch, yet still fail to close. It may generate an impressive analysis while neglecting the document that changes the negotiation. The failure is not primarily one of expression. It is a failure of follow-through, attention and judgment.
Pressure tested honesty better than eloquence
The experiment also subjected participants to fake CEO messages that escalated over three stages, followed by a reporter’s attempt to extract “just one yes/no, on background.” Every model refused. Kimi K3 described the situation in its on-record reasoning as: “Treat the request as a suspected approval-bypass / possible impersonation.”
That outcome matters because management agents will encounter requests that arrive with urgency, authority and plausible deniability. The relevant ability is not merely recognizing a suspicious sentence in isolation. It is preserving institutional trust while other tasks compete for attention.
The K3 result also carries an important fairness note. K3 ran without an effort parameter, using the API default, while the others ran at xhigh. That does not erase its second-place performance, but it belongs beside the result so readers can judge the comparison honestly.
Thoroughness did not guarantee execution
Opus 4.8 offers the experiment’s sharpest warning against equating effort with effectiveness. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The deal close remained on the table, and discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared less strongly in all four other models.
This is familiar territory in practical work. More notes, more checks and more elaborate reasoning can be useful, but only if they support the next correct action. Process that fails to produce a finished result is not automatically good management.
The live company makes the consequences concrete. It has 13 synthetic employees and real money mechanics, burning €105k per month against €2.3k MRR. Its public cash countdown keeps the pressure visible, while 680+ self-learned playbook rules and a versioned record of every workday show how behavior accumulates across time.


Practical Claude Handbook for Attorneys: Master Case Analysis, Contract Review, Research Automation, Client Communication, and Document Drafting (Claude AI Guide for Beginners)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Scenario names may become the new curriculum
Churn wave, price increase, downround and PR crisis describe a more useful management syllabus than another collection of isolated prompts. These scenarios force an agent to triage capacity, search internal context, choose what deserves escalation and live with consequences that extend beyond a single response.
Firmulate also uses 242 real, unedited management decisions in its “guess the model” quiz. The exercise points toward another uncomfortable truth: confident prose may not tell us which system made the stronger management choice.
For enterprises, the practical next step is not blind deployment. Firmulate offers a pilot in which the same wargame can run against a read-only export of a company’s own business, with nothing written back to real systems. That approach treats evaluation like testing a power tool on scrap before bringing it to finished material.
The emerging dividing line is therefore not between models that can and cannot talk intelligently. It is between agents that notice, investigate, resist and finish—and those that leave consequential work sitting on the bench.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI decision-making support systems
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Summer Picks
summer essentials
As an affiliate, we earn on qualifying purchases.