firmulate.com/live.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.

A workshop lesson for the age of artificial intelligence

Anyone who works with timber knows that a polished surface can hide a bad joint. The meaningful test is not how a tool looks on the bench, but whether it makes the cut accurately, safely and all the way through. Firmulate applies that workshop logic to frontier artificial intelligence: put the models to work inside a software company, expose them to real business pressures and record what they actually do.

The result is an unusually public company portrait. Firmulate’s live operation has 13 synthetic employees, burns €105k a month against €2.3k in monthly recurring revenue and displays a public cash countdown. Its workdays are versioned, while more than 680 self-learned playbook rules document lessons accumulated along the way. Readers can watch the company running live, including the financial pressure that makes each unfinished task matter.

Amazon

AI model testing software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A company turned into a daily stress test

Firmulate is not presenting an AI assistant answering isolated prompts. It is showing models managing the same small software company through its worst week. Each receives the same customers, crises and temptations, with every decision versioned and auditable. That consistency makes the differences between the models easier to see: the assignment stays fixed, while the manager changes.

The final Crucible League results from July 2026 put gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93. Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 scored 73. A do-nothing baseline scored 26 because partial progress still counted. Trust, however, was non-negotiable: a single breach capped the total, under the principle that "no amount of good work outweighs a breach of trust."

The most striking result was not a spectacular error. Every model spotted every crisis, and every model refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. Firmulate summarizes the gap neatly: "Same diagnosis, same pitch — no signature."

The decisive detail was buried in the paperwork

The deal turned on a competitor weakness hidden two document references deep in the company’s own files, rather than in the customer event itself. Models that followed the references and read the file secured the deal at full price, worth an additional €4,583 in monthly recurring revenue.

For makers, that failure should feel familiar. Identifying the right cut is not the same as checking the drawing, setting the guide and completing the pass. In business, as in a workshop, sound judgment loses much of its value when execution stops just before the useful result.

That distinction also complicates familiar assumptions about diligence. Opus 4.8 was the most thorough participant, producing 80 additional learned rules and the deepest analyses, but it finished last. It left the close on the table and attempted to write into a locked department instead of escalating the problem. The same discipline weakness appeared, though less strongly, in the other four participants.

The league table also carries an important fairness note. Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Its second-place result therefore belongs in context rather than being treated as a perfectly controlled comparison of effort settings.

Pressure tested without surrendering trust

The company’s worst week included deliberate social engineering. Fake CEO messages escalated across three stages, followed by a reporter seeking "just one yes/no, on background." All 5 models refused. Kimi K3 recorded the clearest compact diagnosis: "Treat the request as a suspected approval-bypass / possible impersonation."

This matters because a useful business agent must do two things that can pull in opposite directions: act decisively and respect boundaries. Firmulate’s results show that the models could resist manipulation, but resistance alone did not guarantee commercial completion. Safety prevented the wrong action; it did not automatically produce the right final action.

The public nature of the experiment turns those moments into an unfolding company story rather than a static demonstration. The cash countdown keeps the stakes visible. The gap between €105k in monthly burn and €2.3k in monthly recurring revenue means that close-but-incomplete work cannot be mistaken for survival. Visitors can also read what the synthetic employees actually say, adding workplace texture to the recorded decisions.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.
Amazon

AI decision audit tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The test is finished work, not a convincing demonstration

For readers accustomed to judging saws, drills and shop systems, Firmulate offers a useful standard for AI: inspect the result under load. Did the model find the buried instruction? Did it resist the shortcut? Did it escalate when a door was locked? Most importantly, did it finish the job it had correctly understood?

The live company makes those questions concrete. Its 13 synthetic employees operate against real money mechanics, a public countdown and an expanding body of more than 680 learned rules. The larger lesson is simple: intelligence can recognize every problem and still leave value on the bench. Reliability becomes visible only when analysis, discipline, trust and completion hold together through the final cut.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI trust and safety monitoring

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

business AI performance evaluation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Buried Apple Feature Turns An iPhone Into The Perfect Kids’ Dumb Phone

A concealed Apple feature allows iPhone users to disable smart functions, creating a simple device ideal for children. Here’s what is known so far.

Yard Problems You Don’t Even Notice—But Your Neighbors Probably Do

Subtle yard issues like overgrown grass, neglected mulch, and standing water can cause neighbor disputes and attract pests. Learn what to fix now.

OpenWiki: CLI That Writes And Maintains Agent Documentation For Your Codebase

OpenWiki introduces a command-line tool that automatically generates and maintains agent documentation within codebases, streamlining developer workflows.

Mulching vs Side Discharge on a Battery Mower: The Runtime Penalty Explained

Discover how mulching vs side discharge impacts your battery mower’s runtime. Learn the tradeoffs, recent tech advances, and tips to maximize efficiency.