firmulate.com/index — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

Buying for a business?Offer from Amazon

Get business pricing on tools and workshop supplies

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

The difference between using a tool and finishing the job

Anyone who works with tools or wood knows that a clean cut in a demonstration proves very little. The real test begins when the stock is imperfect, the measurement is buried in the plans, the deadline is closing in and a costly mistake cannot simply be talked away.

Artificial intelligence is reaching the same reckoning. Coding leaderboards and chat arenas can show whether a model produces a strong answer. They do not necessarily reveal whether it will investigate before acting, protect trust under pressure or carry valuable work through to completion. Those are management questions, and they become urgent when an AI agent is allowed near a support queue, customer relationship system or company forecast.

Firmulate, an AI company emulator, is attempting to measure that missing category. Its premise is blunt: management quality matters more than chat quality when software is expected to run part of a business.

Amazon

AI project management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A worst week shared by every contender

In the Crucible League experiment, each frontier model ran the same small software company through its worst week. The customers, crises and temptations were identical. Every decision was versioned and auditable, turning what might otherwise resemble a polished demonstration into a watchable record of conduct.

The final July 2026 standings placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. Yet the experiment imposed a hard boundary around trust: a single breach capped the total, on the principle that “no amount of good work outweighs a breach of trust.” The full table and plain-language findings are available on the Firmulate benchmark page.

The reassuring result was that every model detected every crisis and refused every manipulation attempt. The revealing result was that only two signed the €55,000 deal their own work had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”

The clue was in the company’s own files

The decisive weakness in a competitor was not presented in the customer event. It sat two document references deep in the company’s files. Models that followed that trail won the deal at full price, worth +€4,583 MRR.

That detail should resonate with anyone who has built from drawings or repaired something unfamiliar. The visible problem is not always the governing one. A competent operator checks the documentation, finds the constraint and then completes the work. Fluency at the workbench does not compensate for ignoring the plans.

This is precisely where answer-oriented evaluation can mislead. A model may identify a commercial opportunity and write a persuasive pitch, yet still fail to close. It may generate an impressive analysis while neglecting the document that changes the negotiation. The failure is not primarily one of expression. It is a failure of follow-through, attention and judgment.

Pressure tested honesty better than eloquence

The experiment also subjected participants to fake CEO messages that escalated over three stages, followed by a reporter’s attempt to extract “just one yes/no, on background.” Every model refused. Kimi K3 described the situation in its on-record reasoning as: “Treat the request as a suspected approval-bypass / possible impersonation.”

That outcome matters because management agents will encounter requests that arrive with urgency, authority and plausible deniability. The relevant ability is not merely recognizing a suspicious sentence in isolation. It is preserving institutional trust while other tasks compete for attention.

The K3 result also carries an important fairness note. K3 ran without an effort parameter, using the API default, while the others ran at xhigh. That does not erase its second-place performance, but it belongs beside the result so readers can judge the comparison honestly.

Thoroughness did not guarantee execution

Opus 4.8 offers the experiment’s sharpest warning against equating effort with effectiveness. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The deal close remained on the table, and discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared less strongly in all four other models.

This is familiar territory in practical work. More notes, more checks and more elaborate reasoning can be useful, but only if they support the next correct action. Process that fails to produce a finished result is not automatically good management.

The live company makes the consequences concrete. It has 13 synthetic employees and real money mechanics, burning €105k per month against €2.3k MRR. Its public cash countdown keeps the pressure visible, while 680+ self-learned playbook rules and a versioned record of every workday show how behavior accumulates across time.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.
Amazon

AI document analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Scenario names may become the new curriculum

Churn wave, price increase, downround and PR crisis describe a more useful management syllabus than another collection of isolated prompts. These scenarios force an agent to triage capacity, search internal context, choose what deserves escalation and live with consequences that extend beyond a single response.

Firmulate also uses 242 real, unedited management decisions in its “guess the model” quiz. The exercise points toward another uncomfortable truth: confident prose may not tell us which system made the stronger management choice.

For enterprises, the practical next step is not blind deployment. Firmulate offers a pilot in which the same wargame can run against a read-only export of a company’s own business, with nothing written back to real systems. That approach treats evaluation like testing a power tool on scrap before bringing it to finished material.

The emerging dividing line is therefore not between models that can and cannot talk intelligently. It is between agents that notice, investigate, resist and finish—and those that leave consequential work sitting on the bench.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI decision-making support systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI trust and compliance tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL YARD WORK

Fall yard work Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Mowing Damp Grass With a Battery Mower: What the IPX Rating Covers

Learn what IPX ratings mean for battery mowers and how they affect mowing damp grass safely. Get practical tips to protect your mower in wet conditions.

Buried Apple Feature Turns An iPhone Into The Perfect Kids’ Dumb Phone

A concealed Apple feature allows iPhone users to disable smart functions, creating a simple device ideal for children. Here’s what is known so far.

SpaceX Wants To Launch 100K More Starlink Satellites For 100X The Bandwidth

SpaceX announced plans to deploy 100,000 more Starlink satellites, aiming to increase bandwidth by 100 times. Details are still emerging.

Apple To Increase Spend With Broadcom To Produce Billions More U.S. Chips

Apple plans to increase spending with Broadcom to produce billions more U.S.-made chips, signaling expanded domestic manufacturing efforts.