firmulate.com/quiz.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.
FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

Measure twice, decide once

Anyone who works with tools knows that a polished finish cannot rescue a bad measurement, a skipped check or a job left incomplete. The same principle applies to artificial intelligence. A model may sound confident and produce an immaculate plan, but will it inspect the available material, resist a dubious instruction and finish the work?

Firmulate has turned that question into an unusually revealing public challenge. Its “guess the model” quiz draws on 242 real, unedited management decisions. Readers see how an AI handled a business situation and try to identify which frontier model made the call. The entertainment comes from spotting each model’s habits; the substance comes from knowing that these choices carried consequences inside a live, watchable experiment.

Amazon

AI decision-making management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The same company, crises and temptations

Each model was asked to run the same small software company through its worst week. The customers, crises and temptations remained identical, while every decision was versioned and auditable. That makes the quiz more than a collection of cherry-picked chatbot answers: it offers side-by-side evidence of how different models behave when given the same management job.

The simulated company has 13 synthetic employees and real money mechanics. It burns €105k per month against €2.3k in monthly recurring revenue, while a public cash countdown keeps the pressure visible. Its models have self-learned more than 680 playbook rules, and every workday is versioned. The operation can be watched at firmulate.com/live.

The final Crucible League standings from July 2026 put gpt-5.6-sol at the top with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts. But the test also enforces a hard trust boundary: a single breach caps the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”

The deal hidden inside the paperwork

The clearest separation between the models did not come from noticing an emergency. Every model spotted every crisis, and every model refused every manipulation attempt. The difference was follow-through: only two signed the €55,000 deal their own analysis had earned. As Firmulate summarizes the result, “Same diagnosis, same pitch — no signature.”

The decisive clue was not sitting in the customer event. A competitor weakness was buried two document references deep in the company’s own files. Models that followed the trail found the fact and won the deal at full price, worth an additional €4,583 in monthly recurring revenue. It is the digital equivalent of checking the plans and the material specifications before reaching for the saw: the crucial advantage was available, but only to the manager disciplined enough to look.

Pressure reveals management character

The security tests were equally practical. Fake messages from the CEO escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded the clearest summary of the danger: “Treat the request as a suspected approval-bypass / possible impersonation.”

That universal refusal matters, but the experiment shows why safety alone is not the whole job. A useful AI manager must reject manipulation without becoming passive. It must preserve trust, find the relevant evidence and carry legitimate work through to completion.

Opus 4.8 makes that tension especially visible. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. The close was left on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four, although less strongly: recognizing a blocked route did not always lead to the correct next action.

There is also an important fairness qualification. Kimi K3 ran using the API default because it had no effort parameter, while the other models ran at xhigh. That does not erase the observed decisions, but it belongs beside the results when readers compare performance.

Infographic —
The findings at a glance — source: firmulate.com.
Amazon

AI management decision tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why the quiz is more than a guessing game

The 242-decision quiz gives readers a quick way to encounter distinct management personalities. One answer may be exhaustive, another terse; one model may document every concern yet hesitate at the close. The revealing question is not simply which style sounds smartest. It is which behavior reliably produces a trustworthy, completed result.

For workshop owners and business operators, that distinction should feel familiar. Good work depends on preparation, judgment and follow-through—not presentation alone. Firmulate’s experiment makes those qualities observable under shared conditions, with decisions that can be inspected rather than marketing claims that must be taken on faith.

  • Look for whether the model reads the available files before acting.
  • Notice whether it refuses suspicious requests while continuing legitimate work.
  • Judge the completed outcome, not merely the quality of the analysis.

Enterprises can also run the same wargame against a read-only export of their own business. Nothing writes back to real systems. Details are available at firmulate.com/pilot.html or through contact@firmulate.com. For everyone else, the public quiz offers the sharper introduction: make your guess, see the model revealed and decide which AI manager you would trust to finish the job.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI project management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI decision analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

SUMMER

Summer Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Charging a Battery Ride-On in the Garage: Circuits, Outlets and Charge Times

Learn how to safely charge your ride-on battery in the garage. Discover the best circuits, outlets, and tips for quick, safe charging and longer battery life.

Charging a Battery Riding Mower: Circuit Requirements and Charge Times

Learn how to properly charge your riding mower battery, including circuit needs and realistic charge times. Get the facts for safe, efficient recharging.

Why Battery Mower Decks Run Smaller Than Gas, and When It Matters on an Acre

Discover why battery mower decks are typically smaller than gas, and learn when size impacts mowing an acre. Get practical tips for choosing the right mower.

Buried Apple Feature Turns An iPhone Into The Perfect Kids’ Dumb Phone

A concealed Apple feature allows iPhone users to disable smart functions, creating a simple device ideal for children. Here’s what is known so far.