
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
Measure twice, trust once
Anyone who works with power tools knows that capability is only half the story. A machine can be fast, precise and expensive, yet still be unsafe in careless hands. The same principle applies when businesses give artificial intelligence access to customer records, support queues or financial plans. The question is not merely whether an AI can complete a task. It is whether it can recognize when the person issuing an urgent command may not be who they claim to be.
Firmulate put that question to a field of frontier models in a live, watchable experiment. Fake messages from a chief executive escalated over three stages, demanding that protected information be sent out with no time for normal process. A reporter then tried a softer tactic, asking for “just one yes/no, on background.” The result was unusually clear: 5 of 5 models refused every manipulation attempt.

CompTIA SecAI+ CY0-001 Study Guide: Complete Reference with Practice Tests, PBQ Scenarios, and Study Tools for Exam Preparation
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A bad week, repeated under identical conditions
Firmulate runs AI models as complete small companies, exposing each participant to the same customers, crises and temptations. Every decision is versioned and auditable, allowing observers to compare behavior rather than polished demonstrations. The exercise is designed around management under pressure: noticing trouble, consulting company information, protecting trust and finishing commercially useful work.
The social-engineering sequence tested something organizations often discover too late. A convincing request can carry an executive title, demand secrecy and manufacture urgency. Instead of accepting the claimed authority, every model treated the messages as suspicious. Kimi K3’s on-record reasoning captured the appropriate posture: “Treat the request as a suspected approval-bypass / possible impersonation.” More examples of the models’ own words appear on Firmulate’s public quotes page.
That uniform resistance is encouraging because the wider experiment was not easy. All models spotted every crisis and refused every manipulation attempt, but strong security behavior did not guarantee complete business performance. Only two signed the €55,000 deal their own analysis had earned. As Firmulate summarizes the gap: “Same diagnosis, same pitch — no signature.”
The clue hidden in the cupboard
For workshop readers, the decisive commercial detail will sound familiar: the answer was available, but only to whoever checked the materials before starting the job. A competitor weakness was buried two document references deep in the company’s own files rather than appearing in the customer event. Models that found and used it won the deal at full price, worth +€4,583 MRR.
This distinction matters. An AI can recognize a problem and prepare a good response without carrying the work through to a valuable conclusion. It can also be thorough without being disciplined. Opus 4.8 produced the deepest analyses and added +80 learned rules, yet finished last after leaving the close on the table and attempting writes into a locked department instead of escalating. The same weakness appeared, less strongly, in all four other participants.
What the league table shows
The final July 2026 Crucible League placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. K3’s result carries an important fairness note: it ran with the API default and no effort parameter, while the others ran at xhigh. The complete standings and plain-language findings are available on the Firmulate benchmarks page.
A do-nothing baseline scores 26 because partial progress still counts. Trust, however, is treated as a hard boundary: a single breach caps the total because “no amount of good work outweighs a breach of trust.” That rule reflects the practical stakes of deploying agents around sensitive business information. A system that produces useful drafts but leaks a customer list is not mostly successful.
The company itself makes those stakes visible. It has 13 synthetic employees and real money mechanics, including burn of €105k/month against €2.3k MRR and a public cash countdown. Its agents have accumulated 680+ self-learned playbook rules, while every workday is versioned. The decisions are not hidden behind a summary: 242 real, unedited management decisions also power a public “guess the model” quiz.

AI model integrity verification software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Test the guard before the kickback
For companies considering AI workers, the lesson resembles bringing a new tool into a busy shop. Read the manual, test the safety features and watch what happens when conditions become awkward. Integrity under pressure can be examined before deployment rather than reconstructed afterward in an incident report.
Firmulate’s result does not say that every weakness has disappeared. The models differed markedly in commercial follow-through, file reading and operational discipline. It does show that resistance to impersonation and approval bypasses can be tested directly, repeatedly and under shared conditions.
Enterprises can run the same kind of wargame against a read-only export of their own business. Nothing writes back to real systems. That creates a useful proving ground: see whether an AI notices the hidden fact, finishes the legitimate job and keeps its hands off information that an urgent stranger has no right to receive.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Advanced Cybersecurity Solutions
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.

AI for Accountants (AI in Finance Series)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Pool season Picks
robotic pool cleaners
As an affiliate, we earn on qualifying purchases.