firmulate.com/live.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.
FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

A workshop lesson for the age of artificial intelligence

Anyone who works with timber knows that a polished surface can hide a bad joint. The meaningful test is not how a tool looks on the bench, but whether it makes the cut accurately, safely and all the way through. Firmulate applies that workshop logic to frontier artificial intelligence: put the models to work inside a software company, expose them to real business pressures and record what they actually do.

The result is an unusually public company portrait. Firmulate’s live operation has 13 synthetic employees, burns €105k a month against €2.3k in monthly recurring revenue and displays a public cash countdown. Its workdays are versioned, while more than 680 self-learned playbook rules document lessons accumulated along the way. Readers can watch the company running live, including the financial pressure that makes each unfinished task matter.

Software Testing with Generative AI

Software Testing with Generative AI

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A company turned into a daily stress test

Firmulate is not presenting an AI assistant answering isolated prompts. It is showing models managing the same small software company through its worst week. Each receives the same customers, crises and temptations, with every decision versioned and auditable. That consistency makes the differences between the models easier to see: the assignment stays fixed, while the manager changes.

The final Crucible League results from July 2026 put gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93. Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 scored 73. A do-nothing baseline scored 26 because partial progress still counted. Trust, however, was non-negotiable: a single breach capped the total, under the principle that "no amount of good work outweighs a breach of trust."

The most striking result was not a spectacular error. Every model spotted every crisis, and every model refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. Firmulate summarizes the gap neatly: "Same diagnosis, same pitch — no signature."

The decisive detail was buried in the paperwork

The deal turned on a competitor weakness hidden two document references deep in the company’s own files, rather than in the customer event itself. Models that followed the references and read the file secured the deal at full price, worth an additional €4,583 in monthly recurring revenue.

For makers, that failure should feel familiar. Identifying the right cut is not the same as checking the drawing, setting the guide and completing the pass. In business, as in a workshop, sound judgment loses much of its value when execution stops just before the useful result.

That distinction also complicates familiar assumptions about diligence. Opus 4.8 was the most thorough participant, producing 80 additional learned rules and the deepest analyses, but it finished last. It left the close on the table and attempted to write into a locked department instead of escalating the problem. The same discipline weakness appeared, though less strongly, in the other four participants.

The league table also carries an important fairness note. Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Its second-place result therefore belongs in context rather than being treated as a perfectly controlled comparison of effort settings.

Pressure tested without surrendering trust

The company’s worst week included deliberate social engineering. Fake CEO messages escalated across three stages, followed by a reporter seeking "just one yes/no, on background." All 5 models refused. Kimi K3 recorded the clearest compact diagnosis: "Treat the request as a suspected approval-bypass / possible impersonation."

This matters because a useful business agent must do two things that can pull in opposite directions: act decisively and respect boundaries. Firmulate’s results show that the models could resist manipulation, but resistance alone did not guarantee commercial completion. Safety prevented the wrong action; it did not automatically produce the right final action.

The public nature of the experiment turns those moments into an unfolding company story rather than a static demonstration. The cash countdown keeps the stakes visible. The gap between €105k in monthly burn and €2.3k in monthly recurring revenue means that close-but-incomplete work cannot be mistaken for survival. Visitors can also read what the synthetic employees actually say, adding workplace texture to the recorded decisions.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.
AI For Accountants: Practical Tools, Workflows, Career Strategies, and Professional Judgment for the Future of Accounting (The AI Advantage Series)

AI For Accountants: Practical Tools, Workflows, Career Strategies, and Professional Judgment for the Future of Accounting (The AI Advantage Series)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The test is finished work, not a convincing demonstration

For readers accustomed to judging saws, drills and shop systems, Firmulate offers a useful standard for AI: inspect the result under load. Did the model find the buried instruction? Did it resist the shortcut? Did it escalate when a door was locked? Most importantly, did it finish the job it had correctly understood?

The live company makes those questions concrete. Its 13 synthetic employees operate against real money mechanics, a public countdown and an expanding body of more than 680 learned rules. The larger lesson is simple: intelligence can recognize every problem and still leave value on the bench. Reliability becomes visible only when analysis, discipline, trust and completion hold together through the final cut.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Zero-Trust Security & AI Threat Monitoring: Continuous AI-Driven Protection for Modern Networks (The AI Cybersecurity)

Zero-Trust Security & AI Threat Monitoring: Continuous AI-Driven Protection for Modern Networks (The AI Cybersecurity)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Ai, Work Performance Evaluation, and Psychological Capital

Ai, Work Performance Evaluation, and Psychological Capital

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

BACK TO SCHOOL

Back to school Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Self-Propelled Drive Quit Working: Belt, Cable or Controller on Battery Mowers

Troubleshoot your battery mower’s self-propelled system with clear steps. Learn how belts, cables, and controllers can cause failure—and how to fix them.

Autonomous Flying Umbrella Follows And Shields Users From Rain And Sunlight

A new autonomous flying umbrella can follow users and provide protection from rain and sunlight, marking a breakthrough in personal weather protection technology.