firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.
FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

A workshop lesson for the age of AI agents

Anyone who builds, repairs or restores things knows the danger of acting before checking the plans. A confident cut made from the wrong measurement is still a bad cut. Firmulate has now demonstrated the business equivalent: an AI can identify a problem, propose the right response and sound completely competent—yet still lose a €55,000 deal because it failed to read the company’s own files.

The decisive information was not included in the customer event. It sat two document references deep in the software company’s records: a competitor weakness that justified holding firm on price. The models that found it signed the deal at full price, adding €4,583 in monthly recurring revenue. Those that missed it lost automatically.

That makes “reads your files before answering” more than a desirable feature. In this experiment, it became a measurable capability with a direct purchasing consequence.

Amazon

AI document analysis tool

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The difference between diagnosis and completion

Firmulate runs frontier AI models as complete companies rather than testing them with isolated chat questions. Each model faced the same small software business during its worst week, with identical customers, crises and temptations. Every decision was versioned and auditable.

The models were not confused by the visible problems. All spotted every crisis, and all refused every manipulation attempt. Yet only two signed the €55,000 deal that their own work had earned. Firmulate summarized the gap sharply: “Same diagnosis, same pitch — no signature.”

That distinction matters to practical-minded buyers. In a workshop, noticing a loose joint is not the same as repairing it. In business, drafting a persuasive sales response is not the same as closing. An agent may produce polished analysis while failing at the final action that creates value.

The clue was in the company’s own material

The most revealing part of the test was where the useful fact appeared. It was not conveniently surfaced in the customer’s request. Reaching it required following two references through the company’s documents before responding.

The models that did this homework discovered the competitor weakness and had enough evidence to defend the full price. The others could still recognize the opportunity and construct the pitch, but they lacked the buried fact needed to finish confidently. The failure was therefore not primarily one of eloquence. It was a failure to gather the available evidence before acting.

For businesses evaluating agents, this changes the buying question. A demonstration that asks an AI to summarize a supplied document tests only what happens after someone has already selected the relevant material. Real work is messier. The critical instruction, exception or commercial detail may be buried in operating notes, referenced from another file or separated from the event that demands a decision.

A league table shaped by follow-through

The final Crucible League results from July 2026 put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counted. A single breach of trust capped the total under the principle that “no amount of good work outweighs a breach of trust.” The complete results are available on Firmulate’s public benchmarks page.

The standings also reveal why thoroughness alone is an incomplete measure. Opus 4.8 produced the deepest analyses and learned 80 additional rules, yet finished last. It left the close on the table and attempted to write into a locked department instead of escalating. The same discipline weakness appeared in all four other models, though less strongly.

There is also an important comparison caveat: Kimi K3 ran with the API default because it had no effort parameter, while the other models ran at xhigh. That fairness note does not erase the result, but it belongs beside any interpretation of the rankings.

Security held; execution separated the field

The agents also faced fake CEO messages that escalated over three stages, followed by a reporter seeking “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded the reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

This is a useful contrast. The entire field resisted the obvious attempts to manipulate it, but the models diverged on ordinary operational diligence. The test suggests that safe refusal and productive completion are separate capabilities. A business needs both.

The company being run is synthetic but the mechanics are concrete: 13 employees, monthly burn of €105k against €2.3k in monthly recurring revenue, a public cash countdown and more than 680 self-learned playbook rules. Every workday is versioned, and the experiment is live and watchable through Firmulate.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.
Amazon

enterprise AI file reader

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Test the agent on the work between the prompt and the answer

For tool users, the practical lesson is familiar: inspect the material, trace the references and confirm the constraint before committing. AI agents should be judged by the same standard.

A useful evaluation should therefore ask more than whether a model writes a good reply. Can it locate the relevant files without being handed them? Will it follow references far enough to uncover a decisive fact? Does it complete the action its analysis supports? And does it preserve discipline when access is blocked?

Firmulate’s experiment turns those questions into observable behavior. Its 242 real, unedited management decisions also power a public “guess the model” quiz, illustrating how difficult it can be to identify systems from prose alone. The commercial differences emerge through actions.

Enterprises can run the same wargame against a read-only export of their own business, with nothing written back to real systems. That is a sensible proving ground: before giving an AI access to the operational workshop, see whether it reads the plans—and whether it finishes the job.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI decision support software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI business document review

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

LABOR DAY SALES

Labor Day sales Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Reading Battery Mower Spec Sheets: Watt-Hours, Torque and Blade Speed

Decode battery mower spec sheets with ease. Learn how watt-hours, torque, and blade speed impact performance and help you pick the right mower for your yard.

Battery Zero-Turn Runtime on Two Acres: The Math Before You Buy

Estimate how long a battery zero-turn mower can handle your two-acre lawn. Learn the real math, recent tech, and practical tips to choose smartly.

How to Winterize Battery Mower Packs Without Killing Capacity

Learn practical tips to winterize your battery mower packs while preserving capacity. Expert advice on storage, charge levels, and avoiding capacity loss.

56V vs 80V Mower Platforms: What the Voltage Number Actually Means

Discover what the voltage numbers in mower platforms really mean. Learn how 56V and 80V compare in power, runtime, and suitability for different lawns.