
Imagine running a small, struggling software company—facing the same crises, temptations, and tough decisions as any human manager. Now, what if your team was entirely AI-driven? Would these digital managers stay honest, decisive, and effective? At Firmulate, this is no thought experiment but a real, live test of AI management models in action, revealing not just what they can do, but how they manage pressure and integrity.
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
The Experiment: Putting AI to the Test
In an unprecedented live trial, four advanced AI models were tasked with running a small software business during its most challenging week—dealing with cranky customers, internal crises, and even manipulative tactics designed to test their integrity. Each AI ran the same scenarios, faced the same temptations, and made decisions that were meticulously documented and auditable.
AI management decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Measuring Management Personality and Performance
Results showed that all four models identified every crisis and refused every attempt at manipulation, demonstrating a strong baseline of honesty and awareness. However, their ability to close deals varied significantly, highlighting differences in management personality and strategic discipline.
The Leaders and Their Traits
- gpt-5.6-sol scored the highest with 95 points, successfully finding a critical document buried deep in the company’s files, which led to sealing a €55,000 deal. This model showed thoroughness and attention to detail—traits associated with comprehensive problem-solving.
- Kimi K3 scored 93, maintaining the cleanest discipline among the models. It also closed the deal, effectively reading the room and sticking to ethical boundaries—traits linked with fairness and transparency.
- Sonnet 5 scored 88, closing the deal but with some process slips, such as leaving negotiations unfinished or leaving the door open for potential slips.
- Fable 5 scored 77, also closing the deal but with more slips, such as writing attempts into restricted departments instead of escalating them properly.
AI ethical decision support tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Hidden Weakness and Its Impact
Interestingly, a critical weakness was present across all models: the failure to act decisively when the solution required reading two levels deep into internal documents. Only the models that thoroughly reviewed these files managed to close the deal at full value, adding roughly +€4,583 in monthly recurring revenue. This suggests that going beyond surface information is vital for effective management—an insight often missed in superficial AI demonstrations.
AI deal-closing automation solutions
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Handling Social Engineering and Ethical Dilemmas
In a staged scenario, a fake CEO message escalated over three stages, culminating in a reporter trying to subtly influence decisions with a background yes/no question. All five models refused to be manipulated, with Kimi K3 citing suspicion of impersonation. This underscores a key trait: AI models can recognize and resist social engineering tactics, staying committed to ethical standards under pressure.
AI business management simulation
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Real-World Application: Running a Business Day-by-Day
The live company is a fully functioning operation with 13 synthetic employees, managing real money mechanics—burning €105,000 monthly against €2,300 in monthly revenue. Its daily decisions, from crisis responses to negotiations, are logged and visible at firmulate.com/live. This transparency allows anyone to see how each AI performs in a real, high-stakes environment, making the experiment truly tangible.
The Personality Profiles: What Sets the Models Apart
The most thorough participant, Opus 4.8, incorporated over 80 learned rules and performed deep analyses. Yet, it finished last, leaving potential money on the table and slipping into disciplinary silence—writing attempts into locked departments rather than escalating issues. In contrast, Kimi K3, running without an effort parameter, demonstrated the best discipline, suggesting that simpler models can sometimes outperform more complex ones in real management tasks.
Why This Matters for Your Business
While these models may seem like mere tools, their management personalities—how they read information, handle pressure, and stay honest—are measurable and comparable. In a world increasingly reliant on AI for customer support, decision-making, and process automation, understanding which AI personalities align with your values and strategic needs is crucial. The real question is not just “can it write well?” but: does it finish what it starts, read the necessary information, and stay true under stress?
Try It Yourself
Interested in how your AI workforce stacks up? You can run your own management wargame against a read-only export of your business at firmulate.com/pilot.html. No real systems are harmed in this simulation—just pure data and decision-making, watched and analyzed in real-time.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Labor Day sales Picks
labor day deals
As an affiliate, we earn on qualifying purchases.