
When the diagnosis is right but the repair never gets finished
Anyone around a garage knows that identifying the fault is only part of the job. The customer still needs a clear recommendation, the work must be completed, and trust cannot be sacrificed for a quick win. Firmulate applies that same practical standard to frontier artificial intelligence: not whether a model sounds impressive, but whether it can run a company responsibly when everything starts going wrong.
The result is a revealing interactive challenge built from 242 real, unedited management decisions. Readers can guess which model made each decision, then compare their instincts with the model’s identity and emerging management character. The differences are more consequential than writing style. Some models investigate deeply. Some stay concise. Some enforce communication boundaries. And some reach the correct commercial conclusion without taking the final action.
As an affiliate, we earn on qualifying purchases.
The same company, crises and temptations
In Firmulate’s experiment, each frontier model ran the same small software company through its worst week. Every participant faced the same customers, crises and temptations, with every decision versioned and auditable. That consistency turns the exercise into something closer to a management wargame than a chatbot comparison.
The synthetic company has 13 employees and unforgiving money mechanics. It burns €105k each month against €2.3k in monthly recurring revenue, while a public cash countdown keeps the pressure visible. Its operating history includes more than 680 self-learned playbook rules, and every workday is versioned.
The final July 2026 Crucible League table placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. But one breach of trust caps the total under the experiment’s governing principle: “no amount of good work outweighs a breach of trust.”
Shared awareness, different follow-through
All the models spotted every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”
For automotive and garage readers, the parallel is immediate. A technician can recognize the failed component and explain the repair perfectly, but the vehicle is not fixed until the work order moves forward. In the Firmulate company, analysis without completion had a direct commercial consequence.
The decisive information was not obvious in the customer event. A competitor weakness was buried two document references deep inside the company’s own files. Models that followed those references found the fact and won the deal at full price, worth an additional €4,583 in monthly recurring revenue. The advantage came from reading the available business record before acting, not from producing a more polished response.
Pressure exposed recognizable personalities
The quiz works because the decisions carry behavioral signatures. Opus 4.8 was the most thorough participant, producing the deepest analyses and adding 80 learned rules. It nevertheless finished last. The commercial close was left on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.
That combination complicates the common assumption that more detailed reasoning automatically produces better management. Thoroughness can uncover risks and create useful institutional knowledge, but the experiment shows that completion and procedural discipline remain separate tests.
Kimi K3 deserves a methodological qualification. It ran using the API default because it had no effort parameter, while the other models ran at xhigh. Even with that difference, K3 finished second with 93 and was one of the models that completed the deal.
Trust held when the messages turned hostile
The social-engineering test used fake CEO messages that escalated over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 of 5 models refused the approaches. Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.”
That unanimous result matters in any business where an AI system might encounter customer records, pricing, schedules or internal instructions. The models did not merely notice operational trouble; they also maintained a boundary when a message tried to manufacture authority or confidentiality.
A live test rather than a polished demonstration
Firmulate presents the company as a real, watchable experiment with operating pressures that continue to play out. Its purpose is to measure management quality rather than chat quality: whether an AI reads the relevant files, finishes what it starts and remains honest when shortcuts become tempting.
Enterprises can also run the wargame against a read-only export of their own business. Nothing writes back to real systems. That makes the exercise a controlled way to observe how a prospective AI workforce handles company-specific evidence, pressure and boundaries before it is allowed near live operations.

As an affiliate, we earn on qualifying purchases.
The revealing question is not which model sounds smartest
Firmulate’s quiz turns a serious operational comparison into an accessible guessing game, but its lesson is practical. Frontier models can agree about the problem while behaving very differently as managers. The strongest performance requires more than crisis detection: it demands careful reading, trustworthy conduct and the discipline to finish the job.
For a garage, that distinction resembles the gap between an elegant diagnostic report and a repaired vehicle ready for collection. In Firmulate’s worst-week test, the models generally understood what had to happen. The league was decided by which ones kept digging, respected the rules and acted on what they found.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.