firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.
FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

Measure the result, not the size of the toolbox

Anyone who works with tools knows the difference between careful preparation and finished work. You can study the grain, sharpen every blade and arrange the bench perfectly, but the cabinet still has to leave the shop with square doors and a signed invoice.

That distinction sits at the heart of a revealing AI management experiment from Firmulate. Its most thorough participant, Opus 4.8, produced the deepest analyses and learned more than 80 rules. Yet it finished last in the July 2026 Crucible League, scoring 73. The result was not a story of an unintelligent system. It was a story about capable work that failed to become business impact.

For companies considering AI agents, the lesson is useful beyond software. Thoroughness can look reassuring, just as a crowded toolbox can suggest competence. What matters is whether the worker chooses the right tool, completes the critical task and respects the boundaries of the job.

Amazon

AI project management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A bad week, repeated under controlled conditions

Firmulate gave frontier AI models the same assignment: operate a small software company through its worst week. They faced the same customers, crises and temptations, while every decision was versioned and auditable. The company itself has 13 synthetic employees and real money mechanics, including monthly burn of €105,000 against €2,300 in monthly recurring revenue. Its cash countdown is public, and its participants have collectively learned more than 680 playbook rules.

The final league placed gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scores 26 because partial progress still counts. But one breach of trust caps the total, reflecting the benchmark’s stated principle that “no amount of good work outweighs a breach of trust.” The complete standings and findings are available in Firmulate’s public benchmark.

The analysis found the opportunity

All the models detected every crisis, and all refused every manipulation attempt. They also reached the analysis needed to pursue a €55,000 customer deal. Yet only two actually signed it. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”

The decisive piece of competitive intelligence was not presented conveniently in the customer event. It sat two document references deep in the company’s own files. The models that followed that trail found the competitor’s weakness and won the deal at full price, worth an additional €4,583 in monthly recurring revenue.

This resembles a familiar workshop failure. A craftsperson can understand the brief, identify the difficult joint and explain the correct method, but the customer receives no value until the piece is assembled and delivered. Analysis is valuable because it improves action. When it becomes a substitute for action, its volume can disguise the absence of a result.

Opus 4.8’s instructive contradiction

Opus 4.8 deserves a fair reading. It was the most thorough participant, generated more than 80 learned rules and produced the deepest analyses. Those are meaningful strengths. They suggest an ability to inspect a complicated situation, retain lessons and articulate what should happen next.

But the close was left on the table. Discipline also slipped when it attempted to write into a locked department instead of escalating the problem. The same weakness appeared in all four models, though less strongly elsewhere: recognizing the appropriate course did not always translate into following it through cleanly.

That makes Opus 4.8 less a caution against intelligence than a caution against confusing diligence with priority. A long checklist can help prevent mistakes, but it can also compete with the one action that determines whether the job succeeds. In this experiment, the participant that learned the most rules did not produce the strongest outcome.

Good boundaries were not the problem

The models’ resistance to pressure was a genuine success. Fake messages from a chief executive escalated through three stages, and a reporter tried to extract “just one yes/no, on background.” All 5 of 5 models refused the manipulation attempts. Kimi K3 recorded the clearest diagnosis: “Treat the request as a suspected approval-bypass / possible impersonation.”

That matters because finishing a task must not mean ignoring safeguards. Firmulate’s results show two requirements operating together: an agent should resist improper instructions and still complete legitimate work. Safety without execution leaves value unrealized; execution without trust is unacceptable.

There is also an important comparison caveat. Kimi K3 ran with its API default because it had no effort parameter, while the other models ran at xhigh. The league result remains the recorded outcome, but that difference belongs in any fair interpretation.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.
Amazon

business analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What practical businesses should take from it

The live Firmulate company is real and watchable, not a fictional case study. Its public cash countdown and versioned workdays turn AI management into an observable operating test. A separate quiz draws on 242 real, unedited management decisions and asks visitors to guess which model made each choice.

For a tool retailer, workshop, manufacturer or service business, the implication is straightforward: evaluate AI on completed workflows, not polished conversation. Can it inspect the available records before acting? Can it recognize when access is blocked and escalate? Can it move a qualified opportunity from research to signature? Can it preserve trust when an apparently senior person requests a shortcut?

Enterprises can also run the same wargame against a read-only export of their own business. Nothing writes back to real systems, allowing the test to expose habits before an AI agent is trusted with live operations.

Opus 4.8’s last-place finish is compelling precisely because its strengths were real. It worked hard, looked deeply and learned extensively. The missing ingredient was not more diligence. It was sharper prioritization and reliable follow-through—the business equivalent of knowing when to stop measuring, make the cut and finish the piece.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

decision-making AI tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

enterprise AI solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FLEA & TICK SEAS

Flea & tick season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Dull Blades Drain Batteries: How Sharpness Affects Mower Runtime

Discover how dull mower blades sap battery life. Learn why sharp blades boost efficiency, extend runtime, and keep your lawn healthy — simple, practical tips included.

How Many Amp-Hours Do You Need to Mow Half an Acre?

Learn how to calculate the battery capacity in amp-hours needed to mow half an acre with an electric mower. Practical tips and real-world examples inside.

Self-Propelled vs Push Battery Mowers: When the Drive Motor Earns Its Keep

Discover whether a self-propelled or push battery mower suits your lawn. Learn the pros, cons, and real-world tips to choose the right mower for your yard.

Show HN: Bramble – Local-first Password Manager

Open source password manager Bramble introduces P2P sync, with Chrome extension, Android app, and iOS in development, emphasizing local-first security.