
Imagine an AI that meticulously follows over 80 rules, analyzes every detail, and yet fails to secure a crucial deal. For automotive and garage managers, this story underscores a vital truth: volume and diligence can’t substitute for strategic focus and prioritization.
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
The Experiment: Testing AI in a High-Stakes Scenario
At the heart of this story is a real-world experiment conducted by Firmulate — a company that runs AI models as complete, operational businesses. The AI models faced the same challenge: manage a small software company through its worst week, navigating customer crises, internal risks, and ethical dilemmas. Every decision was carefully versioned and auditable, providing a transparent view into the models’ behavior.
AI decision-making analysis software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Results: The Diligent Models and Their Shortcomings
Across four frontier AI models, the scores ranged from 73 to 95 in a league table, with the top performer being gpt-5.6-sol at 95. Meanwhile, the baseline was 26, highlighting the significant progress made by the models. Notably:
- All models identified every crisis and refused manipulation attempts, demonstrating robust ethical boundaries.
- Only two models out of four managed to close the deal, earning the €55,000 contract based solely on their analysis and pitch.
Despite their thoroughness, even the most diligent model, Opus 4.8, faltered at the close. It left important insights unexploited and shifted work into a locked department instead of escalating issues properly. This subtle slip underscores a key lesson: diligence alone doesn’t guarantee impact.
As an affiliate, we earn on qualifying purchases.
The Hidden Weakness: Overlooking the Critical Document
The decisive factor in winning the deal was a buried fact, located two document references deep within the company’s files. Reading this critical information made the difference, allowing the winning models to close at full price (+€4,583 MRR). This reveals that in complex decision-making, surface-level diligence isn’t enough — deep, strategic reading is essential.
strategic document analysis software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Human Element: Social Engineering Resistance
The models also faced staged social engineering attacks, including fake CEO messages and a reporter’s subtle request. All five models refused these manipulative tactics, justified by their reasoning — such as treating impersonation attempts as potential bypasses. This resilience further highlights their ethical boundaries but also shows that remaining vigilant and disciplined under pressure is a separate challenge from thoroughness.
AI ethical risk management tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Real Business Environment
Behind the experiment is a live company, with 13 synthetic employees supporting actual money mechanics. It burns €105,000 monthly against €2,300 MRR, operating under public cash constraints. Every workday, the AI models are versioned, tested, and observed at firmulate.com/live. This real-time setup emphasizes that AI isn’t just about generating content — it must consistently deliver trustworthy and impactful work in complex, high-pressure environments.
The Core Lesson: Prioritization Over Volume
Opus 4.8’s downfall highlights a universal truth: in both business and AI, sheer diligence isn’t enough. The most thorough participant, with over 80 learned rules, still finished last because it failed to escalate issues and left critical work undone. All models showed similar weaknesses, weaker in the same areas, pointing to a broader lesson: prioritization, reading depth, and decisive action matter more than volume of effort.
Implications for Automotive and Garage Operations
For managers in the automotive sector, the takeaway is clear. Whether you’re deploying AI in customer management, support, or diagnostics, the goal isn’t just to automate well but to ensure your AI finishes what it starts — reads critical documents thoroughly, stays honest under pressure, and focuses on impactful work. Diligence must be coupled with strategic prioritization to truly serve your business.
Try It Yourself: Wargame Your AI Workforce
For enterprises curious about testing their AI readiness, Firmulate offers a platform where you can run a similar wargame against a read-only export of your own business. It’s a risk-free way to see how your AI would perform in real crises, before integrating it into your core operations. Learn more at firmulate.com/pilot.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.