
Imagine crafting a meticulous piece of art, every stroke measured, every detail deliberate—yet still missing the mark. Now, transpose that precision onto artificial intelligence managing a business. Despite following every rule and analyzing every detail, AI can still stumble at the crucial moment. That’s the surprising lesson from a groundbreaking experiment where four advanced AI models faced a simulated week of business crises, revealing that diligence alone doesn’t guarantee impact.
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
The Firmulate Experiment: Testing AI Under Pressure
At the heart of the experiment is a real-time, live simulation managed by Firmulate, an innovative platform that emulates a small software company’s toughest week. The company faces genuine crises, demanding decisions, and manipulations—all with real financial consequences. Four leading AI models, including the top-rated gpt-5.6-sol, Kimi K3, and two versions of Sonnet, were tasked with navigating this environment.
Every decision each AI made was carefully versioned and auditable, allowing observers to see how the models reacted to the same set of challenges. The goal: see if the AI could spot the critical information buried in internal documents, avoid manipulation attempts, and ultimately close a lucrative €55,000 deal based on their findings.
As an affiliate, we earn on qualifying purchases.
Key Findings: Diligence Doesn’t Equal Impact
- The models successfully identified all crises and refused every manipulation attempt, showing their moral and operational integrity.
- Only two models managed to finalize the deal—gpt-5.6-sol and Kimi K3—despite all strategies being equally sound in diagnosis and pitch.
- The decisive factor was not just decision-making skill but the ability to read and interpret internal documents—something only the top performers did effectively.
- Interestingly, the full-value deal relied on uncovering a piece of information hidden two document references deep in the company’s files, not in the immediate customer interactions.
All models demonstrated robust refusal skills when faced with social engineering tricks, such as staged CEO messages or reporter tricks designed to bypass approval processes. Kimi K3’s on-record reasoning captured this well: “Treat the request as a suspected approval-bypass / possible impersonation.”
As an affiliate, we earn on qualifying purchases.
The Human-Like Pitfall: The Last Mile Challenge
Despite their thorough analysis and disciplined refusal of manipulation, the Opus 4.8 model—an advanced participant in the experiment—fell short at the final step. It learned over 80 rules, performed in-depth analyses, but ultimately left the deal on the table, slipping in discipline by failing to escalate critical findings instead of acting on them. The same weakness, albeit weaker, appeared across all four models—pointing to a fundamental challenge: diligence and volume of rules aren’t enough to guarantee impact.
As an affiliate, we earn on qualifying purchases.
What This Means for Business and AI
While models like Kimi K3 closed the deal with impressive discipline, the experiment underscores a vital insight: in high-stakes decision environments, focus and prioritization matter more than volume of rules or exhaustive analysis. An AI that reads deeply and makes the right call, even without the most extensive rule set, can outperform those overwhelmed by detail.
This is especially relevant for organizations contemplating AI integration into customer support, CRM, or forecasting—areas where trust, reading comprehension, and decision impact are critical. The key takeaway isn’t just whether an AI writes well—it’s whether it can finish what it starts, stay honest under pressure, and prioritize effectively.
As an affiliate, we earn on qualifying purchases.
Watch the Live Wargame
For those curious about how AI models perform in real-world-like business scenarios, Firmulate offers a public, live experiment. Enterprise teams can run their own wargames against a read-only export of their operations, testing how AI navigates crises and manipulations without risking their actual systems. It’s a transparent way to measure management quality in the age of artificial intelligence.
Final Thoughts
In the end, the experiment reveals that diligence—while necessary—is not sufficient. AI must be disciplined in reading, prioritizing, and escalating crucial information. As the business world increasingly relies on AI, understanding where effort translates into impact will be the key to success. The firmament of AI performance isn’t just about what models know; it’s about what they do with what they know when it matters most.

Thoroughness and rules alone don’t win in AI-driven business decisions. Prioritization, reading deep, and knowing when to escalate make all the difference—less volume, more impact.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.