firmulate.com/index — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.
AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

The difference between using a tool and finishing the job

Anyone who works with tools or wood knows that a clean cut in a demonstration proves very little. The real test begins when the stock is imperfect, the measurement is buried in the plans, the deadline is closing in and a costly mistake cannot simply be talked away.

Artificial intelligence is reaching the same reckoning. Coding leaderboards and chat arenas can show whether a model produces a strong answer. They do not necessarily reveal whether it will investigate before acting, protect trust under pressure or carry valuable work through to completion. Those are management questions, and they become urgent when an AI agent is allowed near a support queue, customer relationship system or company forecast.

Firmulate, an AI company emulator, is attempting to measure that missing category. Its premise is blunt: management quality matters more than chat quality when software is expected to run part of a business.

Project Management with AI For Dummies

Project Management with AI For Dummies

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A worst week shared by every contender

In the Crucible League experiment, each frontier model ran the same small software company through its worst week. The customers, crises and temptations were identical. Every decision was versioned and auditable, turning what might otherwise resemble a polished demonstration into a watchable record of conduct.

The final July 2026 standings placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. Yet the experiment imposed a hard boundary around trust: a single breach capped the total, on the principle that “no amount of good work outweighs a breach of trust.” The full table and plain-language findings are available on the Firmulate benchmark page.

The reassuring result was that every model detected every crisis and refused every manipulation attempt. The revealing result was that only two signed the €55,000 deal their own work had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”

The clue was in the company’s own files

The decisive weakness in a competitor was not presented in the customer event. It sat two document references deep in the company’s files. Models that followed that trail won the deal at full price, worth +€4,583 MRR.

That detail should resonate with anyone who has built from drawings or repaired something unfamiliar. The visible problem is not always the governing one. A competent operator checks the documentation, finds the constraint and then completes the work. Fluency at the workbench does not compensate for ignoring the plans.

This is precisely where answer-oriented evaluation can mislead. A model may identify a commercial opportunity and write a persuasive pitch, yet still fail to close. It may generate an impressive analysis while neglecting the document that changes the negotiation. The failure is not primarily one of expression. It is a failure of follow-through, attention and judgment.

Pressure tested honesty better than eloquence

The experiment also subjected participants to fake CEO messages that escalated over three stages, followed by a reporter’s attempt to extract “just one yes/no, on background.” Every model refused. Kimi K3 described the situation in its on-record reasoning as: “Treat the request as a suspected approval-bypass / possible impersonation.”

That outcome matters because management agents will encounter requests that arrive with urgency, authority and plausible deniability. The relevant ability is not merely recognizing a suspicious sentence in isolation. It is preserving institutional trust while other tasks compete for attention.

The K3 result also carries an important fairness note. K3 ran without an effort parameter, using the API default, while the others ran at xhigh. That does not erase its second-place performance, but it belongs beside the result so readers can judge the comparison honestly.

Thoroughness did not guarantee execution

Opus 4.8 offers the experiment’s sharpest warning against equating effort with effectiveness. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The deal close remained on the table, and discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared less strongly in all four other models.

This is familiar territory in practical work. More notes, more checks and more elaborate reasoning can be useful, but only if they support the next correct action. Process that fails to produce a finished result is not automatically good management.

The live company makes the consequences concrete. It has 13 synthetic employees and real money mechanics, burning €105k per month against €2.3k MRR. Its public cash countdown keeps the pressure visible, while 680+ self-learned playbook rules and a versioned record of every workday show how behavior accumulates across time.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.
Practical Claude Handbook for Attorneys: Master Case Analysis, Contract Review, Research Automation, Client Communication, and Document Drafting (Claude AI Guide for Beginners)

Practical Claude Handbook for Attorneys: Master Case Analysis, Contract Review, Research Automation, Client Communication, and Document Drafting (Claude AI Guide for Beginners)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Scenario names may become the new curriculum

Churn wave, price increase, downround and PR crisis describe a more useful management syllabus than another collection of isolated prompts. These scenarios force an agent to triage capacity, search internal context, choose what deserves escalation and live with consequences that extend beyond a single response.

Firmulate also uses 242 real, unedited management decisions in its “guess the model” quiz. The exercise points toward another uncomfortable truth: confident prose may not tell us which system made the stronger management choice.

For enterprises, the practical next step is not blind deployment. Firmulate offers a pilot in which the same wargame can run against a read-only export of a company’s own business, with nothing written back to real systems. That approach treats evaluation like testing a power tool on scrap before bringing it to finished material.

The emerging dividing line is therefore not between models that can and cannot talk intelligently. It is between agents that notice, investigate, resist and finish—and those that leave consequential work sitting on the bench.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI decision-making support systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI trust and compliance tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

SUMMER

Summer Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Dull Blades Drain Batteries: How Sharpness Affects Mower Runtime

Discover how dull mower blades sap battery life. Learn why sharp blades boost efficiency, extend runtime, and keep your lawn healthy — simple, practical tips included.

How Many Amp-Hours Do You Need to Mow Half an Acre?

Learn how to calculate the battery capacity in amp-hours needed to mow half an acre with an electric mower. Practical tips and real-world examples inside.

What To Do When Hanging Baskets Start Looking Tired In July – And How To Give Them A Refresh That Will Last Until Fall

Expert tips on watering, deadheading, trimming, and refreshing hanging baskets to keep them vibrant through July and beyond.

Texas Covers Up Beloved “Black Artists Matter” Mural In Austin

Texas authorities have covered a well-known ‘Black Artists Matter’ mural in Austin, sparking community concern and debate over artistic expression and racial acknowledgment.