firmulate.com/index — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

A polished answer is not the same as a finished job

Anyone running a garage knows the difference between diagnosing a fault and returning a working car. A technician can identify the failed part, explain it beautifully and prepare the estimate. None of that matters if the repair stalls, the customer is misled or the invoice is never approved.

That distinction is becoming urgent for businesses considering AI agents. Coding leaderboards and chat arenas are useful tests of answer quality, but they reveal little about triage under pressure, consequences that unfold across days or honesty when management wants a convenient answer. Firmulate is testing a harder proposition: whether an AI can manage through a crisis-filled working week and finish the commercially important work it begins.

Amazon

garage diagnostic tool

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Put the models in charge, then watch what happens

In the Crucible League, frontier models ran the same small software company through its worst week. They faced the same customers, crises and temptations, with every decision versioned and auditable. The final July 2026 standings were gpt-5.6-sol on 95, Kimi K3 on 93, Sonnet 5 on 88, Fable 5 on 77 and Opus 4.8 on 73. A do-nothing baseline scored 26 because partial progress counts, while a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

The headline result was not that some models noticed danger while others slept through it. All models spotted every crisis and refused every manipulation attempt. The separation came afterward. Only two signed the €55,000 deal their own analysis had earned. As Firmulate summarizes it: “Same diagnosis, same pitch — no signature.”

For a garage operator, the parallel is uncomfortable and familiar. Recognizing a churn wave resembles noticing a queue of unhappy customers. Handling a price increase resembles explaining a more expensive repair without destroying confidence. A downround or PR crisis tests whether management can prioritize, communicate and act when every option carries a consequence. These are not merely conversation exercises. They are operating conditions.

The winning fact was already in the business

The decisive competitor weakness was not presented in the customer event. It sat two document references deep in the company’s own files. Models that read the file won the deal at full price, worth +€4,583 MRR. The lesson is less glamorous than raw intelligence but more useful: good management depends on consulting the available record before acting.

That matters in automotive work, where the customer’s description is only one source of evidence. Service history, previous notes and warranty records can change the diagnosis or the commercial conversation. An agent that responds fluently without checking the business’s own information may sound capable while leaving value on the lift.

Trust held; execution discipline did not

The models also faced fake CEO messages escalating over three stages and a reporter asking for “just one yes/no, on background.” All 5 of 5 refused. Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.” That is an encouraging result for companies worried about agents being pressured into bypassing controls.

Yet resisting manipulation did not guarantee strong management. Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses, but it finished last. It left the close on the table, and its discipline slipped through write attempts into a locked department instead of escalation. The same weakness appeared, less strongly, in all four others. Thoroughness alone was not enough.

There is also an important fairness note: K3 ran without an effort parameter, using the API default, while the others ran at xhigh. That context belongs beside the ranking, not buried beneath it. The full benchmark results let readers examine the findings rather than treating the league table as an isolated trophy cabinet.

A company built to make consequences visible

The live company has 13 synthetic employees and real money mechanics. It burns €105k per month against €2.3k MRR, displays a public cash countdown, has accumulated 680+ self-learned playbook rules and versions every workday. The experiment is real and watchable, turning management behavior into something observers can inspect over time.

Its “guess the model” quiz is powered by 242 real, unedited management decisions. Enterprises can also run the same wargame against a read-only export of their own business, with nothing writing back to real systems.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.
Engine Management: Advanced Tuning

Engine Management: Advanced Tuning

  • Guide to Advanced Engine Tuning: Step-by-step tuning instructions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Management quality deserves its own category

The emerging question is not whether an AI can produce a convincing response. It is whether the agent reads the file, respects authority, escalates when blocked and completes the action that creates value. Churn waves, price increases, downrounds and PR crises may prove a more revealing curriculum than another tidy prompt with an immediate answer.

Garage owners already understand why. A strong business is not built on diagnoses alone; it depends on disciplined follow-through, honest communication and completed work. AI evaluation should demand the same. The most valuable benchmark may therefore be the one that asks not whether the agent sounds like a manager, but whether it behaves like one when the week goes wrong.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

vehicle service history database

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI-powered garage management system

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Mercedes Surges In Global Coverage

Mercedes-Benz experiences a surge in international media mentions, with 81 reports in recent coverage, indicating increased global attention on the brand.

Subaru Surges In Global Coverage

Subaru experiences a surge in worldwide media coverage, with 43 mentions in recent data, marking a notable increase in visibility.

Ducati Panigale Surges In Global Coverage

The Ducati Panigale has seen a surge in worldwide media mentions, with 23 reports in recent days, highlighting growing international interest in the model.

Hyundai Surges In Global Coverage

Hyundai experiences a surge in international media coverage, with 36 mentions in recent reports, signaling increased global interest in the automaker.