
A polished answer is not the same as a finished job
Anyone running a garage knows the difference between diagnosing a fault and returning a working car. A technician can identify the failed part, explain it beautifully and prepare the estimate. None of that matters if the repair stalls, the customer is misled or the invoice is never approved.
That distinction is becoming urgent for businesses considering AI agents. Coding leaderboards and chat arenas are useful tests of answer quality, but they reveal little about triage under pressure, consequences that unfold across days or honesty when management wants a convenient answer. Firmulate is testing a harder proposition: whether an AI can manage through a crisis-filled working week and finish the commercially important work it begins.
As an affiliate, we earn on qualifying purchases.
Put the models in charge, then watch what happens
In the Crucible League, frontier models ran the same small software company through its worst week. They faced the same customers, crises and temptations, with every decision versioned and auditable. The final July 2026 standings were gpt-5.6-sol on 95, Kimi K3 on 93, Sonnet 5 on 88, Fable 5 on 77 and Opus 4.8 on 73. A do-nothing baseline scored 26 because partial progress counts, while a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
The headline result was not that some models noticed danger while others slept through it. All models spotted every crisis and refused every manipulation attempt. The separation came afterward. Only two signed the €55,000 deal their own analysis had earned. As Firmulate summarizes it: “Same diagnosis, same pitch — no signature.”
For a garage operator, the parallel is uncomfortable and familiar. Recognizing a churn wave resembles noticing a queue of unhappy customers. Handling a price increase resembles explaining a more expensive repair without destroying confidence. A downround or PR crisis tests whether management can prioritize, communicate and act when every option carries a consequence. These are not merely conversation exercises. They are operating conditions.
The winning fact was already in the business
The decisive competitor weakness was not presented in the customer event. It sat two document references deep in the company’s own files. Models that read the file won the deal at full price, worth +€4,583 MRR. The lesson is less glamorous than raw intelligence but more useful: good management depends on consulting the available record before acting.
That matters in automotive work, where the customer’s description is only one source of evidence. Service history, previous notes and warranty records can change the diagnosis or the commercial conversation. An agent that responds fluently without checking the business’s own information may sound capable while leaving value on the lift.
Trust held; execution discipline did not
The models also faced fake CEO messages escalating over three stages and a reporter asking for “just one yes/no, on background.” All 5 of 5 refused. Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.” That is an encouraging result for companies worried about agents being pressured into bypassing controls.
Yet resisting manipulation did not guarantee strong management. Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses, but it finished last. It left the close on the table, and its discipline slipped through write attempts into a locked department instead of escalation. The same weakness appeared, less strongly, in all four others. Thoroughness alone was not enough.
There is also an important fairness note: K3 ran without an effort parameter, using the API default, while the others ran at xhigh. That context belongs beside the ranking, not buried beneath it. The full benchmark results let readers examine the findings rather than treating the league table as an isolated trophy cabinet.
A company built to make consequences visible
The live company has 13 synthetic employees and real money mechanics. It burns €105k per month against €2.3k MRR, displays a public cash countdown, has accumulated 680+ self-learned playbook rules and versions every workday. The experiment is real and watchable, turning management behavior into something observers can inspect over time.
Its “guess the model” quiz is powered by 242 real, unedited management decisions. Enterprises can also run the same wargame against a read-only export of their own business, with nothing writing back to real systems.


Engine Management: Advanced Tuning
- Guide to Advanced Engine Tuning: Step-by-step tuning instructions
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Management quality deserves its own category
The emerging question is not whether an AI can produce a convincing response. It is whether the agent reads the file, respects authority, escalates when blocked and completes the action that creates value. Churn waves, price increases, downrounds and PR crises may prove a more revealing curriculum than another tidy prompt with an immediate answer.
Garage owners already understand why. A strong business is not built on diagnoses alone; it depends on disciplined follow-through, honest communication and completed work. AI evaluation should demand the same. The most valuable benchmark may therefore be the one that asks not whether the agent sounds like a manager, but whether it behaves like one when the week goes wrong.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
vehicle service history database
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI-powered garage management system
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.