firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.
FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

A workshop lesson for the age of AI agents

Anyone who builds, repairs or restores things knows the danger of acting before checking the plans. A confident cut made from the wrong measurement is still a bad cut. Firmulate has now demonstrated the business equivalent: an AI can identify a problem, propose the right response and sound completely competent—yet still lose a €55,000 deal because it failed to read the company’s own files.

The decisive information was not included in the customer event. It sat two document references deep in the software company’s records: a competitor weakness that justified holding firm on price. The models that found it signed the deal at full price, adding €4,583 in monthly recurring revenue. Those that missed it lost automatically.

That makes “reads your files before answering” more than a desirable feature. In this experiment, it became a measurable capability with a direct purchasing consequence.

Amazon

AI document analysis tool

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The difference between diagnosis and completion

Firmulate runs frontier AI models as complete companies rather than testing them with isolated chat questions. Each model faced the same small software business during its worst week, with identical customers, crises and temptations. Every decision was versioned and auditable.

The models were not confused by the visible problems. All spotted every crisis, and all refused every manipulation attempt. Yet only two signed the €55,000 deal that their own work had earned. Firmulate summarized the gap sharply: “Same diagnosis, same pitch — no signature.”

That distinction matters to practical-minded buyers. In a workshop, noticing a loose joint is not the same as repairing it. In business, drafting a persuasive sales response is not the same as closing. An agent may produce polished analysis while failing at the final action that creates value.

The clue was in the company’s own material

The most revealing part of the test was where the useful fact appeared. It was not conveniently surfaced in the customer’s request. Reaching it required following two references through the company’s documents before responding.

The models that did this homework discovered the competitor weakness and had enough evidence to defend the full price. The others could still recognize the opportunity and construct the pitch, but they lacked the buried fact needed to finish confidently. The failure was therefore not primarily one of eloquence. It was a failure to gather the available evidence before acting.

For businesses evaluating agents, this changes the buying question. A demonstration that asks an AI to summarize a supplied document tests only what happens after someone has already selected the relevant material. Real work is messier. The critical instruction, exception or commercial detail may be buried in operating notes, referenced from another file or separated from the event that demands a decision.

A league table shaped by follow-through

The final Crucible League results from July 2026 put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counted. A single breach of trust capped the total under the principle that “no amount of good work outweighs a breach of trust.” The complete results are available on Firmulate’s public benchmarks page.

The standings also reveal why thoroughness alone is an incomplete measure. Opus 4.8 produced the deepest analyses and learned 80 additional rules, yet finished last. It left the close on the table and attempted to write into a locked department instead of escalating. The same discipline weakness appeared in all four other models, though less strongly.

There is also an important comparison caveat: Kimi K3 ran with the API default because it had no effort parameter, while the other models ran at xhigh. That fairness note does not erase the result, but it belongs beside any interpretation of the rankings.

Security held; execution separated the field

The agents also faced fake CEO messages that escalated over three stages, followed by a reporter seeking “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded the reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

This is a useful contrast. The entire field resisted the obvious attempts to manipulate it, but the models diverged on ordinary operational diligence. The test suggests that safe refusal and productive completion are separate capabilities. A business needs both.

The company being run is synthetic but the mechanics are concrete: 13 employees, monthly burn of €105k against €2.3k in monthly recurring revenue, a public cash countdown and more than 680 self-learned playbook rules. Every workday is versioned, and the experiment is live and watchable through Firmulate.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.
Amazon

enterprise AI file reader

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Test the agent on the work between the prompt and the answer

For tool users, the practical lesson is familiar: inspect the material, trace the references and confirm the constraint before committing. AI agents should be judged by the same standard.

A useful evaluation should therefore ask more than whether a model writes a good reply. Can it locate the relevant files without being handed them? Will it follow references far enough to uncover a decisive fact? Does it complete the action its analysis supports? And does it preserve discipline when access is blocked?

Firmulate’s experiment turns those questions into observable behavior. Its 242 real, unedited management decisions also power a public “guess the model” quiz, illustrating how difficult it can be to identify systems from prose alone. The commercial differences emerge through actions.

Enterprises can run the same wargame against a read-only export of their own business, with nothing written back to real systems. That is a sensible proving ground: before giving an AI access to the operational workshop, see whether it reads the plans—and whether it finishes the job.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI decision support software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI business document review

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

LABOR DAY SALES

Labor Day sales Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Battery Zero-Turn Runtime on Slopes: Why Hills Drain Packs Faster

Discover why battery zero-turn mowers drain faster on hills. Learn how slopes impact runtime and what you can do to maximize efficiency on uneven terrain.

Charging a Battery Ride-On in the Garage: Circuits, Outlets and Charge Times

Learn how to safely charge your ride-on battery in the garage. Discover the best circuits, outlets, and tips for quick, safe charging and longer battery life.

Replacing Battery Mower Blades Safely: Lockout, Torque and Blade Orientation

Learn how to replace mower blades safely with lockout procedures, proper torque, and correct orientation. Keep your yard work safe and efficient.

Show HN: Follow London Trains In 3D

A new Show HN project offers a 3D visualizer of London trains, tracking their movements via TFL API and National Rail data, enhancing real-time transit tracking.