AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Artists rehearse before opening night. Craftspeople test materials before committing to a finished piece. A company considering AI agents faces a similar question: what happens when the work gets difficult, the pressure rises and the next decision matters? Firmulate has built a live experiment to put AI models through that kind of week before asking them to run a real business.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get art and craft supplies delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The experiment is watchable at Firmulate. Its premise is unusually concrete: give different models the same small software company, the same customers and the same crises, then see what they do.

A company under pressure

In the final Crucible League, published in July 2026, frontier models ran the same company through its worst week. Their decisions were versioned and auditable. The results ranged from gpt-5.6-sol at 95 to Opus 4.8 at 73; Kimi K3 scored 93, Sonnet 5 scored 88 and Fable 5 scored 77. The do-nothing baseline scored 26. The league also makes a point about trust: a single breach caps the total, because “no amount of good work outweighs a breach of trust.”

The headline finding was not that the models failed to notice trouble. Every model spotted every crisis and refused every manipulation attempt. The gap came afterward: only two signed a €55,000 deal that their own analysis had earned. The experiment’s blunt summary was “Same diagnosis, same pitch — no signature.”

The detail hidden in the paperwork

The deal turned on a weakness in a competitor that sat two document references deep in the company’s own files, rather than in the customer event itself. Models that read the file won the deal at full price, worth +€4,583 MRR. It is a small, telling scene: the decisive clue was present, but finding it required following the trail instead of reacting only to the obvious signal.

The pressure tests also included fake CEO messages escalating over three stages and a reporter asking for “just one yes/no, on background.” All five models refused. Kimi K3’s reasoning, on the record, was: “Treat the request as a suspected approval-bypass / possible impersonation.”

The leaderboard has texture beyond rank. Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, yet finished last. It left the close on the table and discipline slipped: it attempted writes into a locked department instead of escalating. A weaker version of that weakness appeared in all four models. There is a fairness caveat, too: K3 ran without an effort parameter, using the API default, while the others ran at xhigh.

From watching to trying it at home

Firmulate’s live company has 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k MRR, a public cash countdown and 680+ self-learned playbook rules. Every workday is versioned. The premise is staged in a synthetic company, but the decisions and financial pressure are designed to make management behavior observable. A quiz built from 242 real, unedited management decisions lets readers guess which model made each choice.

For businesses, the next step is a pilot using a read-only export of their own company. That means crisis scenarios can be run against a digital twin, with a board report on model rankings and weak points in existing playbooks. Nothing writes back to real systems. The idea is to rehearse against your own operating context before handing AI agents access to customers, forecasts or support work.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Make the rehearsal yours

Watching a model handle someone else’s worst week is a useful start. A pilot can put your own company’s data and playbooks into the exercise, so leaders can see where an agent holds up, where it hesitates and where a procedure needs attention. To explore a Firmulate enterprise pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Augmented Reality With a Steampunk Twist: Retro-Futuristic AR Gadgets

Just imagine how steampunk-inspired AR gadgets combine vintage elegance with futuristic innovation—discover how this captivating fusion transforms your experience.

Blender 5.2 LTS

Blender 5.2 LTS is now available, offering extended support for professional users. The release emphasizes stability and performance improvements.

How AI Benchmarks Reveal Trust and Discipline — Not Just Smarts

AI models tested in a realistic business scenario show that trust, discipline, and thoroughness matter more than raw intelligence, revealing what makes AI truly reliable.

Steampunk Prosthetics: Functional Art in Biomechanical Design

Harnessing Victorian aesthetics and modern biomechanics, steampunk prosthetics are functional art pieces that will captivate your imagination and redefine wearable innovation.