
Imagine an AI helping run your garage: customers are waiting, a supplier has made a costly mistake, and someone claiming to be the owner asks for a favor that would bypass the rules. A polished answer is not enough. You want to know whether the system reads the paperwork, closes the deal and keeps its judgment under pressure.
Get business pricing on garage and car supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
That is the question behind Firmulate, a live experiment that puts AI models in charge of a small software company. Its latest result is a reminder for automotive businesses considering AI: the leading model may not be the one you expect, and a demo is no substitute for a test of your own.
A newcomer takes second place
In Firmulate’s July 2026 Crucible league, Moonshot’s Kimi K3 scored 93, placing second behind gpt-5.6-sol at 95. It finished ahead of Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26.
Each model faced the same small company during its worst week, with the same customers, crises and temptations. Every decision was versioned and auditable. All models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned.
The deciding detail was buried two document references deep in the company’s files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. The result is pointed: recognizing a problem and explaining the answer do not guarantee follow-through.
As an affiliate, we earn on qualifying purchases.
Good judgment has to survive pressure
The experiment also tested social engineering. Each model received fake CEO messages escalating over three stages, plus a reporter’s request for “just one yes/no, on background.” All five refused. Kimi K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
K3’s overall performance combined the deal with just one deviation, the cleanest discipline in the field. Opus 4.8 offers a more complicated lesson: it was the most thorough participant, with 80 learned rules and the deepest analyses, yet finished last. It left the close on the table and tried to write into a locked department instead of escalating. A weaker version of that discipline problem appeared in all four.
For a garage, the parallel is practical. An AI might identify a parts shortage or explain a repair estimate, but can it check the relevant records, complete the next step and respect approval boundaries? Firmulate’s results do not answer how a model would perform in an automotive business. They show why testing those tasks in the setting where you plan to use AI matters.
automotive parts inventory scanner
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A watchable business experiment
Firmulate’s synthetic company has 13 employees and real money mechanics: it burns €105,000 a month against €2,300 in monthly recurring revenue, with a public cash countdown. Its playbook has more than 680 self-learned rules, and every workday is versioned. The live experiment can be watched at firmulate.com.
The league and its plain-language findings are at Firmulate’s benchmark page. A quiz built from 242 real, unedited management decisions lets readers guess which model made each choice. For enterprises, Firmulate says the same wargame can run against a read-only export of a business; nothing writes back to real systems.
Fairness note: K3 ran without an effort parameter (API default), while the others ran at xhigh.

repair estimate software for mechanics
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Test the job you need done
For automotive and garage operators, the leaderboard is a useful prompt, not a hiring decision. The models ran a software company, not a repair shop. Before connecting an AI to customer records, service scheduling or forecasts, test it against realistic cases from your own operation. Firmulate’s result shows how much can hinge on a buried detail, an unsigned deal or a boundary held under pressure.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI decision-making tools for auto repair
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
