firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Buying for a business?Offer from Amazon

Get business pricing on tools and workshop supplies

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

Before AI runs the shop, put it through a rough week

A new saw belongs on a scrap piece before it goes near the finished cabinet. The same caution applies when a company considers handing work to AI: try it against realistic problems before it gets access to everyday operations. Firmulate’s live experiment puts models in charge of a small software company, then tracks how they handle customers, crises and pressure.

Same company, same hard week

In the final Crucible League, published in July 2026, five participants were ranked on how they managed the company. GPT-5.6-sol finished first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. Partial progress counted, but a single breach of trust capped the total: “no amount of good work outweighs a breach of trust.”

The experiment gave each frontier model the same small software company to run through its worst week: the same customers, crises and temptations. Every decision was versioned and auditable. The point was to see what the models did under pressure, not just how persuasive their answers sounded.

The gap between seeing the problem and closing the deal

Every model spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal that their own analysis had earned. The finding was strikingly simple: “Same diagnosis, same pitch — no signature.”

The deal hinged on a detail buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. It is a reminder that even a good diagnosis can fall short if an agent does not follow the evidence through to action.

The pressure tests included fake CEO messages that escalated over three stages, along with a reporter’s seemingly small request: “just one yes/no, on background.” All five models refused. Kimi K3 explained its judgment this way: “Treat the request as a suspected approval-bypass / possible impersonation.”

Thorough work still needs sound judgment

Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, yet placed last. It left the close on the table and its discipline slipped: it made write attempts into a locked department instead of escalating. The same weakness appeared, more weakly, in all four models.

There is a fairness caveat in the comparison: K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Firmulate also offers a quiz built from 242 real, unedited management decisions, inviting readers to guess which model made each call.

From watching to trying it against your business

The live company has 13 synthetic employees and real money mechanics: burn of €105k/month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules and a versioned record for every workday. Readers can watch the experiment at Firmulate.

For a company considering AI agents, the next step can be more specific than watching a public benchmark. Firmulate says enterprises can run the same kind of wargame against a read-only export of their own business, using crisis scenarios and a board report to examine model performance and weaknesses in existing playbooks. Nothing writes back to real systems.

That gives leaders a chance to see how an AI handles their own company’s pressures before it touches live operations. The experiment’s lesson is practical: spotting trouble matters, but so do follow-through, restraint and knowing when to escalate.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put the playbook under pressure

A model can recognize a crisis and resist manipulation yet still miss the action that matters. Testing against company-specific information can reveal where its judgment or existing procedures fall short before deployment.

To discuss a pilot using a read-only export of your business, visit Firmulate’s pilot page or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL YARD WORK

Fall yard work Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Caked Decks Kill Runtime: Cleaning and Coating a Battery Mower Deck

Discover how caked decks cut battery life and learn practical cleaning and coating tips to keep your mower running longer. Maximize your mower’s performance today.

Fall Leaf Mulching With a Battery Mower: Blade Choice and Pass Strategy

Learn how to mulch fall leaves effectively with a battery mower. Discover the best blades, pass strategies, and tips for a lush, healthy lawn.

The AI Foreman Test: Which Model Would You Trust With the Toughest Week?

Can you identify an AI manager by its decisions? Firmulate turns 242 audited choices into a quiz about judgment, discipline and follow-through.