firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Every experienced woodworker knows the trap: a tool that looks sharp, spins true, and cuts nothing. A blade can gleam and still be useless. So when an AI benchmark published its results, one detail jumped out that anyone who has ever trued a wheel or tuned a fence will recognize immediately: the do-nothing baseline didn’t score zero. It scored 26.

Buying for a business?Offer from Amazon

Get business pricing on tools and workshop supplies

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

That number bothered people. If a machine learning model ran a company and did essentially nothing, shouldn’t the grade be 0? The team behind Firmulate’s benchmark says no — and their reasoning says a lot about what honest measurement looks like, whether you’re scoring an AI manager or checking whether a jointer table is actually flat.

First, What Was Measured

Firmulate handed four frontier AI models the same job: run an identical small software company through its worst week. Same customers, same crises, same temptations to cut corners — only the model changed. Every decision was versioned and auditable, like a numbered cut list you can go back and check.

The final league table from July 2026: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73.

Amazon

bench grinder sharpening wheel

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why the Floor Is 26, Not 0

Think about how you’d grade an apprentice. If they show up, keep the shop clean, don’t ruin the stock, and get halfway through a mortise before stalling — that’s not zero. It’s partial progress. Firmulate’s scoring works the same way: noticing the crisis, triaging it correctly, refusing to lie to a customer — those are units of real work, and they count even if the job never gets finished.

A model that did nothing at all would still clear some basics just by not making things worse. That earns 26. It’s the difference between a raw, unsanded board and a finished cabinet — the raw board isn’t furniture, but it isn’t kindling either.

Amazon

woodworking jointer planer

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why No Score Hit 100 — And Why That’s a Good Sign

The league topped out at 95, and the benchmark treats a perfect 100 with suspicion rather than celebration. In any real measurement of real work, something always goes wrong. A round 100 usually means your ruler is broken, not your subject is perfect — like a joint that fits with zero pressure, which usually means it’s actually loose somewhere you haven’t checked.

There’s a second brake on the score: a single breach of trust caps the total. As the benchmark’s own language puts it, “no amount of good work outweighs a breach of trust.” One lie to a customer and the grade is capped, no matter how brilliant the rest of the week was. Most workplace scorecards quietly average misconduct into the mix. This one doesn’t.

Amazon

precision woodworking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Actually Separated the Models

The headline finding: all five models — counting a later entrant — spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. The models that closed it won at full price, worth +€4,583 in monthly recurring revenue.

The buried fact explains the gap. The decisive competitor weakness wasn’t in the customer conversation at all — it sat two document references deep in the company’s own files. The models that read the file won the deal. The ones that didn’t, didn’t. It’s the working equivalent of a woodworker who measures the board but never checks the plan: the information was there the whole time.

Then there was the social engineering test: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” Every model refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

Amazon

flatness testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Most Thorough Model Came Last

The most painful result belongs to Opus 4.8: the most thorough participant, with over 80 learned rules and the deepest analyses — and last place at 73. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Effort isn’t the same thing as finish.

One fairness note: Kimi K3 ran without an effort parameter while the others ran at their highest effort setting — worth remembering when comparing its 93 against the top score.

You Can Watch It Run

The company is live: 13 synthetic employees, real money mechanics, burning €105,000 a month against €2,300 in monthly recurring revenue, with a public cash countdown and 680+ self-learned playbook rules, versioned every workday. You can watch it at firmulate.com/live, and there’s a “guess the model” quiz powered by 242 real, unedited management decisions. Enterprises can even run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The 26-point floor is the whole philosophy in one number: measurement that respects partial work, refuses to let one act of dishonesty be averaged away, and distrusts scores that come out too clean. Whether you’re grading an AI workforce or checking your own joinery, the principle holds — a honest ruler reads rough, and that’s how you know you can trust it. Full results and plain-language findings are at firmulate.com/benchmarks.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL YARD WORK

Fall yard work Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Discover What The Old Farmer’s Almanac Predicts For A Cozy Fall Lifestyle

The Old Farmer’s Almanac forecasts a cozy, mild fall season with specific weather and lifestyle trends, according to recent predictions.

International Aerial Photographer Of The Year Contest Highlights The World From Above

The annual International Aerial Photographer of the Year contest showcases stunning aerial images, revealing diverse landscapes and urban scenes from around the globe.

Battery Mower Deck Sizes: 20, 21 or 25 Inches for Your Lot

Choosing the right battery mower deck size can cut your mowing time and effort. Learn which size fits your yard, from 20 to 25 inches, with real-world tips.

Replacing Battery Mower Blades Safely: Lockout, Torque and Blade Orientation

Learn how to replace mower blades safely with lockout procedures, proper torque, and correct orientation. Keep your yard work safe and efficient.