
Imagine managing an ice cream shop during the busiest week of the year. Do you trust your assistant to read the fine print, resist shortcuts, and keep your reputation intact? Now, what if your assistant was an AI, trained with different personalities and decision styles? Welcome to the world of AI management experiments, where real-world decisions are tested under pressure—and the results might surprise you.
The Live AI Company Wargame: Real Decisions, Real Consequences
At the heart of this experiment is a live AI-powered business simulation that runs every weekday at firmulate.com. It mimics a small software firm with 13 synthetic employees, real financial mechanics, and a public cash countdown—much like a bakery rushing to meet demand during a holiday rush. The goal? Measure how well AI models handle crises, temptations, and ethical dilemmas, not just their chat skills.
AI decision-making tools for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Models and Their Personalities
Four frontier AI models faced the same challenging week, each with a distinct decision style. Their scores from the latest Crucible League competition tell the story:
- gpt-5.6-sol scored 95 and was the only one to find the hidden data in the company’s files, sealing a €55,000 deal.
- Kimi K3 scored 93, closed the same deal, and showed the cleanest discipline—refusing manipulative requests.
- Sonnet 5 scored 88, also closed the deal, but with a few process slips.
- Fable 5 scored 77, again closing, but less consistently.
AI ethics and compliance software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Crises, Manipulations, and Honesty
Every model faced identical crises—customer complaints, safety breaches, and ethical dilemmas. Remarkably, all four identified every crisis and refused every attempt at manipulation, including staged social engineering scenarios involving fake CEO messages and reporter tricks. For example, all models rejected a background request asking for a quick approval, citing suspicion of impersonation—a key sign of ethical discipline.
AI document reading and comprehension tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Hidden Weakness: Trust and Document Reading
The real differentiator came down to a buried detail in the company’s files. The models that proactively read and understood these references secured the full deal at +€4,583 monthly recurring revenue. Those that didn’t missed the crucial info and left money on the table, illustrating how deep document comprehension impacts real business outcomes.
As an affiliate, we earn on qualifying purchases.
Personality Profiles and Decision Styles
One model, Opus 4.8, stood out for its thorough analysis—learning over 80 rules and performing the deepest assessments. Yet, it finished last in the league because it struggled with closing discipline, leaving opportunities unclaimed and slipping into departmental silos instead of escalating issues. Meanwhile, Kimi K3, without an effort parameter, maintained the highest fairness and discipline, closing deals efficiently and ethically.
Implications for Business and AI
This experiment underscores an essential point: the question isn’t whether AI can write well or generate convincing chatter. It’s whether AI can finish what it starts, stay honest under pressure, and read the critical documents before acting. For companies considering AI integration into CRM, support, or forecasting, these behavioral qualities are just as vital as raw performance metrics.
Try It Yourself
If you’re curious about how your own business might fare, you can run the same wargame against a read-only export of your operations. It’s a safe way to test your AI’s decision consistency before deployment—without risking real money or customer trust. Visit firmulate.com/pilot.html to learn more and get started.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html