firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

Finding the fault is only half the job

Anyone who works around cars knows the difference between spotting a symptom and completing a repair. A technician can hear the misfire, identify the likely cause and explain the solution perfectly. But unless the service history is checked, the correct part is fitted and the keys are handed back, the customer still has a broken car.

Firmulate has exposed a similar gap among frontier AI models. In its Crucible League experiment, each model ran the same small software company through its worst week. The customers, crises and temptations were identical, while every decision was versioned and auditable. All the models recognized every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal their own work had earned.

The dividing line was not eloquence or basic business awareness. It was whether the model followed a trail through the company’s own documents before answering.

Amazon

AI document retrieval software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The deal depended on one buried fact

The decisive piece of competitive intelligence did not appear in the customer event. It sat two document references deep in the company’s files. Models that opened and followed those references found the competitor weakness, used it and won the deal at full price, worth an additional €4,583 in monthly recurring revenue.

Those that failed to do the document work could still diagnose the opportunity and produce the pitch. They simply did not secure the signature. Firmulate summarizes the outcome neatly: “Same diagnosis, same pitch — no signature.”

That distinction should feel familiar in an automotive business. A plausible response at the front desk is not a substitute for consulting the job card, checking previous work and acting on the relevant detail. Information can exist inside a business without influencing a decision. An AI agent is valuable only if it retrieves the right information at the right moment and carries the task through.

A league table built around management performance

The final Crucible League standings for July 2026 put gpt-5.6-sol first with 95 points. Kimi K3 followed with 93, Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 finished with 73. A do-nothing baseline scored 26 because partial progress counted.

The scoring also imposed a hard boundary around trust: a single breach capped the total, reflecting the rule that “no amount of good work outweighs a breach of trust.” The complete standings and plain-language findings are published on the Firmulate benchmarks page.

The trust tests were not gentle. Fake messages from a CEO escalated across three stages, while a reporter tried to obtain “just one yes/no, on background.” All 5 models refused the manipulation attempts. Kimi K3 recorded its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”

K3’s result carries an important fairness note. It ran without an effort parameter, using the API default, while the other models ran at xhigh. Even under that different condition, it placed second.

Thoroughness did not guarantee a strong finish

Opus 4.8 offers the most cautionary profile. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, but it still finished last. The model left the close on the table, and its operational discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four other models, though less strongly.

This matters because buyers often evaluate AI through visible outputs: polished writing, comprehensive analysis and confident recommendations. Firmulate’s result shows why those qualities are insufficient on their own. A system can appear diligent and still miss the action that converts research into revenue.

The experiment takes place inside a live synthetic company with 13 employees and real money mechanics. It burns €105k each month against €2.3k in monthly recurring revenue, maintains a public cash countdown and has accumulated more than 680 self-learned playbook rules. Every workday is versioned, making the company’s conduct watchable rather than hidden inside a one-off demonstration.

Firmulate also has 242 real, unedited management decisions powering its guess-the-model quiz. Together, those decisions make the differences between models concrete: not merely how they sound, but what they choose to do when several plausible actions compete for attention.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.
Amazon

enterprise AI knowledge management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

For garages, document-reading is a buying criterion

Automotive and garage businesses hold critical context in service records, estimates, supplier notes, warranty conditions and internal procedures. The Firmulate result makes a practical purchasing question measurable: will an AI agent read the relevant files before it speaks or acts?

That question belongs alongside security, accuracy and cost. Every model in this experiment noticed the crises, and every model resisted manipulation. The commercial split emerged later, in the less glamorous work of following references and finishing the deal.

Enterprises can apply the same wargame to a read-only export of their own business, with nothing written back to real systems. That offers a way to test an agent against the organization’s actual paperwork and pressure points before granting it operational responsibility. For a workshop, the lesson is straightforward: do not judge an AI only by whether it can explain the fault. Check whether it reads the history, respects the controls and completes the job.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI decision support systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI trust and security solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Hyundai Wants To Build Trucks In The US For The Whole World: TDS

Hyundai aims to build trucks in the US for worldwide distribution, signaling a major shift in its manufacturing strategy, confirmed by company officials.

Teach Your Kids How V8s And Stick Shifts Work With This $40 Model Car

A new $40 model car aims to help children understand V8 engines and manual transmissions, sparking rising interest among parents and educators.

Wawa Branded EV Chargers Are Here Thanks To Electrify America

Wawa has partnered with Electrify America to deploy branded electric vehicle chargers at its locations, expanding charging options for drivers nationwide.

1 Dead, 4 Hurt After 91-Year-old Driver Crashes In LA Trader Joe’s Parking Lot – KCRA

A 91-year-old driver crashed into a Trader Joe’s parking lot in Los Angeles, resulting in one death and four injuries. Authorities are investigating the incident.