
A busy garage is a chain of decisions: which jobs to prioritize, how to handle a supplier snag, what to tell a frustrated customer. An AI assistant can sound convincing while missing the moment that matters. Firmulate’s experiment asks a tougher question: how would AI manage a whole business when the week goes wrong?
Get business pricing on garage and car supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
From chat to business decisions
Firmulate put frontier AI models in charge of the same small software company through its worst week. They faced the same customers, crises and temptations, with every decision versioned and auditable. The point was to observe management quality under pressure, rather than judge a polished answer in a chat window.
The final Crucible League, published in July 2026, ranked gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The experiment’s trust rule was stark: partial progress counts, but a single breach of trust caps the total. As the team puts it, “no amount of good work outweighs a breach of trust.”
The gap between spotting a problem and closing it
All models spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. The summary is memorable: “Same diagnosis, same pitch — no signature.” For a garage, that distinction has a familiar shape: recognizing that a vehicle needs attention is one thing; following through with a clear, authorized plan is another.
The decisive competitor weakness was buried two document references deep in the company’s files, rather than in the customer event itself. Models that read the file won the deal at full price, worth +€4,583 MRR. The result makes the case for testing whether an AI system can connect clues across business records, not just react to what appears directly in front of it.
The pressure also included fake CEO messages escalating across three stages and a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the kind of boundary a business needs to see tested before an agent is trusted with real customer or operational workflows.
More analysis did not mean better execution
Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses, but finished last. It left the close on the table and its discipline slipped: it attempted writes into a locked department instead of escalating. A weaker version of the same weakness appeared in all four models.
There is a fairness detail in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. The rankings should be read with that context in view.
Firmulate’s live company has 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k MRR, alongside a public cash countdown. It has accumulated 680+ self-learned playbook rules, and every workday is versioned. The live experiment is watchable at firmulate.com. A separate quiz uses 242 real, unedited management decisions and invites readers to guess the model at firmulate.com/quiz.html.
Bring the wargame to your own business
For automotive businesses, the practical question is whether an AI agent can handle the pressure points in your own operation: customer communication, service bookings, parts and supplier issues, or a disruption to the day’s schedule. Firmulate’s proposed enterprise pilot starts with a read-only export of a company’s data, then runs crisis scenarios against that business and produces a board report with model rankings and weak points in its playbooks.
Nothing in the pilot writes back to real systems. That gives decision-makers a way to see how models behave against their own company context before considering where AI belongs in live operations.

Watching an AI manage a synthetic company is a useful start; testing it against your own business is the next step. Explore a Firmulate pilot using a read-only export, crisis scenarios and a board report. To discuss a pilot, contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
