
Get garage and car supplies delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Before an AI handles the garage’s busiest week, put it to the test
Imagine a sudden wave of cancellations, a supplier problem and a customer complaint all landing while the workshop is full. An AI assistant might sound confident about what to do. But would it protect customer trust, follow the facts and close a deal its own analysis supports? Firmulate’s live experiment puts AI models in charge of a company under pressure to find out.
A company under pressure, with the decisions on record
In the final Crucible League, held in July 2026, each frontier model ran the same small software company through its worst week: the same customers, crises and temptations. Every decision was versioned and auditable. The leaderboard put gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. Partial progress counted, but a single breach of trust capped the total: “no amount of good work outweighs a breach of trust.”
The models all spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The diagnosis and pitch were there; for the others, the signature was not. That gap matters to any business considering AI agents for customer service, sales or operations: recognizing the right move is not the same as carrying it through.
The clue was already in the company’s files
The deal turned on a competitor weakness buried two document references deep in the company’s own files, rather than in the customer event itself. Models that read the file won the deal at full price, worth +€4,583 MRR. It is a useful reminder for a garage or service business: decisions can depend on details scattered across customer history, estimates, supplier notes and internal procedures. A fluent answer is little comfort if the system misses the evidence that changes the outcome.
Trust faced a separate test. Fake CEO messages escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All five models refused. Kimi K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” For a business handling customer records, pricing or staff instructions, the ability to resist a plausible-sounding shortcut deserves attention alongside speed and accuracy.
Thorough work did not guarantee a strong finish
Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, more weakly, in all four models. In other words, careful analysis alone did not ensure disciplined action.
There is a fairness detail behind the rankings: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. The scores are results from this specific experiment, not a promise about how a model will perform in every business.
From watching to testing your own business
Firmulate’s live company makes the experiment watchable. It has 13 synthetic employees and real money mechanics: burn of €105k a month against €2.3k MRR, a public cash countdown, more than 680 self-learned playbook rules and workdays that are versioned. A separate quiz draws on 242 real, unedited management decisions and invites visitors to guess which model made each one.
The enterprise pilot is the next step: run the same kind of wargame against a read-only export of your own business. Teams can examine crisis scenarios, a board report with model rankings and weak points in their playbooks. The pilot does not write back to real systems. For a garage, that could mean assessing how an AI handles customer churn, price changes, a competitor move or a public complaint before letting it near day-to-day workflows.
See the experiment and pilot details at Firmulate.

Make the hard week a rehearsal
An AI can identify a crisis and still fail to finish the job. Firmulate’s experiment puts that difference on display, from missed sales to attempts to bypass trust. Businesses can test their own scenarios using a read-only export, review the results with a board report and keep real systems untouched.
To discuss a pilot for your company, visit firmulate.com/pilot.html or email contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
