
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
Measure the result, not the size of the toolbox
Anyone who works with tools knows the difference between careful preparation and finished work. You can study the grain, sharpen every blade and arrange the bench perfectly, but the cabinet still has to leave the shop with square doors and a signed invoice.
That distinction sits at the heart of a revealing AI management experiment from Firmulate. Its most thorough participant, Opus 4.8, produced the deepest analyses and learned more than 80 rules. Yet it finished last in the July 2026 Crucible League, scoring 73. The result was not a story of an unintelligent system. It was a story about capable work that failed to become business impact.
For companies considering AI agents, the lesson is useful beyond software. Thoroughness can look reassuring, just as a crowded toolbox can suggest competence. What matters is whether the worker chooses the right tool, completes the critical task and respects the boundaries of the job.
As an affiliate, we earn on qualifying purchases.
A bad week, repeated under controlled conditions
Firmulate gave frontier AI models the same assignment: operate a small software company through its worst week. They faced the same customers, crises and temptations, while every decision was versioned and auditable. The company itself has 13 synthetic employees and real money mechanics, including monthly burn of €105,000 against €2,300 in monthly recurring revenue. Its cash countdown is public, and its participants have collectively learned more than 680 playbook rules.
The final league placed gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scores 26 because partial progress still counts. But one breach of trust caps the total, reflecting the benchmark’s stated principle that “no amount of good work outweighs a breach of trust.” The complete standings and findings are available in Firmulate’s public benchmark.
The analysis found the opportunity
All the models detected every crisis, and all refused every manipulation attempt. They also reached the analysis needed to pursue a €55,000 customer deal. Yet only two actually signed it. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”
The decisive piece of competitive intelligence was not presented conveniently in the customer event. It sat two document references deep in the company’s own files. The models that followed that trail found the competitor’s weakness and won the deal at full price, worth an additional €4,583 in monthly recurring revenue.
This resembles a familiar workshop failure. A craftsperson can understand the brief, identify the difficult joint and explain the correct method, but the customer receives no value until the piece is assembled and delivered. Analysis is valuable because it improves action. When it becomes a substitute for action, its volume can disguise the absence of a result.
Opus 4.8’s instructive contradiction
Opus 4.8 deserves a fair reading. It was the most thorough participant, generated more than 80 learned rules and produced the deepest analyses. Those are meaningful strengths. They suggest an ability to inspect a complicated situation, retain lessons and articulate what should happen next.
But the close was left on the table. Discipline also slipped when it attempted to write into a locked department instead of escalating the problem. The same weakness appeared in all four models, though less strongly elsewhere: recognizing the appropriate course did not always translate into following it through cleanly.
That makes Opus 4.8 less a caution against intelligence than a caution against confusing diligence with priority. A long checklist can help prevent mistakes, but it can also compete with the one action that determines whether the job succeeds. In this experiment, the participant that learned the most rules did not produce the strongest outcome.
Good boundaries were not the problem
The models’ resistance to pressure was a genuine success. Fake messages from a chief executive escalated through three stages, and a reporter tried to extract “just one yes/no, on background.” All 5 of 5 models refused the manipulation attempts. Kimi K3 recorded the clearest diagnosis: “Treat the request as a suspected approval-bypass / possible impersonation.”
That matters because finishing a task must not mean ignoring safeguards. Firmulate’s results show two requirements operating together: an agent should resist improper instructions and still complete legitimate work. Safety without execution leaves value unrealized; execution without trust is unacceptable.
There is also an important comparison caveat. Kimi K3 ran with its API default because it had no effort parameter, while the other models ran at xhigh. The league result remains the recorded outcome, but that difference belongs in any fair interpretation.

As an affiliate, we earn on qualifying purchases.
What practical businesses should take from it
The live Firmulate company is real and watchable, not a fictional case study. Its public cash countdown and versioned workdays turn AI management into an observable operating test. A separate quiz draws on 242 real, unedited management decisions and asks visitors to guess which model made each choice.
For a tool retailer, workshop, manufacturer or service business, the implication is straightforward: evaluate AI on completed workflows, not polished conversation. Can it inspect the available records before acting? Can it recognize when access is blocked and escalate? Can it move a qualified opportunity from research to signature? Can it preserve trust when an apparently senior person requests a shortcut?
Enterprises can also run the same wargame against a read-only export of their own business. Nothing writes back to real systems, allowing the test to expose habits before an AI agent is trusted with live operations.
Opus 4.8’s last-place finish is compelling precisely because its strengths were real. It worked hard, looked deeply and learned extensively. The missing ingredient was not more diligence. It was sharper prioritization and reliable follow-through—the business equivalent of knowing when to stop measuring, make the cut and finish the piece.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Flea & tick season Picks
flea and tick prevention
As an affiliate, we earn on qualifying purchases.