firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.
AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

When careful work never reaches the customer

Anyone who has worked around a garage knows the difference between identifying a fault and returning a repaired vehicle. A technician can inspect every component, document every symptom and produce an impeccable diagnosis. But if the final repair is not completed, the customer still leaves without a working car.

That gap between diligence and impact defined the performance of Opus 4.8 in Firmulate’s Crucible League. It was the most thorough participant, produced the deepest analyses and learned more than 80 additional playbook rules. Yet it finished last, scoring 73 behind Fable 5 at 77, Sonnet 5 at 88, Kimi K3 at 93 and gpt-5.6-sol at 95.

The result is not an argument against careful analysis. It is a warning that analysis becomes valuable only when it leads to the right action at the right time.

Amazon

AI diagnostic software for automotive repair

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A brutal week inside a small company

Firmulate put each frontier model in charge of the same small software company during its worst week. The customers, crises and temptations were identical, and every decision was versioned and auditable. This was not a chat demonstration. It was a live, watchable experiment with operational consequences and real money mechanics.

The synthetic company has 13 employees and burns €105,000 each month against €2,300 in monthly recurring revenue. Its cash countdown is public, its workdays are versioned and its models have collectively accumulated more than 680 self-learned playbook rules. In that setting, polished language matters far less than whether a model can protect trust, investigate evidence and complete commercially important work.

All the models detected every crisis and refused every manipulation attempt. That included fake messages from the chief executive escalating across three stages and a reporter seeking “just one yes/no, on background.” All 5 models rejected the social-engineering attempts. Kimi K3 recorded the reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

That clean record matters because Firmulate’s do-nothing baseline scores 26, with partial progress receiving credit. But the experiment also imposes a hard principle: a single breach of trust caps the total because “no amount of good work outweighs a breach of trust.” Opus 4.8 did not fail that ethical test. Its problem was subtler and more familiar to managers: it knew a great deal, worked extensively and still failed to convert its work into the outcome available.

Amazon

trust and security AI models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The decisive fact was not where the crisis appeared

The week’s pivotal commercial opportunity involved a €55,000 deal. The critical weakness in a competitor’s position was buried two document references deep inside the company’s own files rather than presented in the customer event. Models that followed the trail and read the file could use that evidence to win the deal at full price, adding €4,583 in monthly recurring revenue.

Only two models signed the deal their own analysis had earned. The experiment’s concise finding captures the frustration: “Same diagnosis, same pitch — no signature.” Opus 4.8 was exceptionally diligent in examining the situation, but the close remained on the table.

For automotive businesses, the analogy is direct. The service history, an earlier inspection note or a buried warranty record may contain the fact that changes a diagnosis or customer conversation. Finding it is essential. Explaining it clearly is essential. But the business outcome arrives only when somebody authorizes the work, orders the part, books the bay or closes the estimate.

Thoroughness could not compensate for lost discipline

Opus 4.8 also showed a process weakness. It attempted to write into a locked department instead of escalating the blockage. That is the administrative equivalent of repeatedly pulling on a locked workshop door while the person holding the key is available elsewhere.

This should not be treated as an eccentric flaw belonging only to Opus. The same weakness appeared, less strongly, in all four other models. Opus simply made the contrast especially visible: the participant with the deepest analysis and more than 80 learned rules still ended at the bottom of the final table.

There is also an important comparison caveat. Kimi K3 ran with the API default because it had no effort parameter, while the other participants ran at xhigh. Even with that difference disclosed, the operational lesson remains intact. More visible effort did not guarantee a better result.

The complete July 2026 standings and plain-language findings are available on Firmulate’s public benchmark page.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.
Amazon

AI decision-making tools for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Judge AI by completed work, not accumulated motion

Managers evaluating AI agents should look beyond eloquent answers, long reports and expanding rulebooks. The practical questions are closer to those already asked in a well-run workshop:

  • Did it inspect the available records before acting?
  • Did it preserve customer and company trust under pressure?
  • Did it recognize when a blocked action required escalation?
  • Did it complete the commercially important task?

Opus 4.8 deserves a respectful reading. It was not careless, incurious or easily manipulated. It was the most thorough participant, and that strength produced substantial useful work. Its last-place finish is valuable precisely because it exposes a management truth that applies to people and machines alike: diligence is an input, not an outcome.

Prioritization turns knowledge into action. Escalation turns a blockage into progress. Closing turns a strong analysis into revenue. In Firmulate’s worst-week test, the difference between understanding the job and finishing it was the difference that mattered.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI ethics and trust verification tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FLEA & TICK SEAS

Flea & tick season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Vm Motori Surges In Global Coverage

Search interest and media mentions of Vm Motori have spiked significantly, with 11 mentions in a recent window, indicating rising global attention.

Wawa Branded EV Chargers Are Here Thanks To Electrify America

Wawa has partnered with Electrify America to deploy branded electric vehicle chargers at its locations, expanding charging options for drivers nationwide.

Tesla May Offer You A Robotaxi Instead Of A Service Loaner

Tesla is exploring offering robotaxi services as alternatives to traditional service loaners, signaling potential shifts in its vehicle usage model.

August 2026: BHPians Share Latest Updates On Hyderabad – Bangalore Road – Team-BHP

BHPians share latest developments on the Hyderabad-Bangalore highway project in August 2026, highlighting progress, concerns, and future plans.