
Get baking supplies delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
What Ice Cream and AI Management Have in Common?
Just as choosing the right ingredients makes your signature ice cream stand out, selecting the best AI model can determine whether your business thrives or dives. Recent experiments show that behind the scenes of AI decision-making are crucial qualities that matter — honesty, discipline, and the ability to finish what it starts. And in a real-world test, the newcomer among AI models managed to beat some of the most established players, all without any special effort parameter.
As an affiliate, we earn on qualifying purchases.
The Real-World AI Company Experiment
At Firmulate, a live and watchable experiment pits four leading AI models against one another—each tasked with running a small, real software company through its toughest week. This is no simple demo; it’s a fully operational test where every decision and action is tracked, versioned, and auditable. The goal? To see which AI can manage crises, resist manipulation, and ultimately close profitable deals, just like a human executive.
The League Table of Performance
- gpt-5.6-sol scored 95, leading the pack by finding buried information in the company files and closing the deal.
- The newcomer Kimi K3 scored 93, narrowly behind, demonstrating the cleanest discipline of the field and winning the same deal — at full price.
- Sonnet 5 scored 88, also closing the deal but with a few process slips.
- Fable 5 scored 77, while Opus 4.8 finished with 73, both closing the deal despite a few weaknesses.
Remarkably, all models identified every crisis and refused manipulative tactics, such as fake CEO messages or reporter tricks. Yet, only Kimi K3 and gpt-5.6-sol actually signed the €55,000 deal their own analysis justified—showing the importance of disciplined decision-making.
The Hidden Weakness in the Files
The decisive advantage for Kimi K3 and gpt-5.6-sol was their ability to read and analyze company files to uncover buried information that wasn’t obvious from the customer’s immediate story. This subtle but crucial insight made the difference between just closing a deal and closing it at full value—adding €4,583 in monthly recurring revenue.
Discipline Under Pressure
The experiment also tested resilience against sleight-of-hand tactics. All models refused to accept fake approvals or background requests to bypass protocol, with Kimi K3 explicitly reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This disciplined stance is vital in real business, where manipulative tactics are common.
The Real Business Is Live and Losing Money
The company managed by these models is not a simulation; it’s a live operation with 13 synthetic employees, burning €105,000 monthly against €2,300 in monthly revenue. Every workday, the decision-making processes are versioned and transparent, providing a unique window into how AI can run real companies. Watch this ongoing experiment at firmulate.com/live.
As an affiliate, we earn on qualifying purchases.
The Surprising Result: A Newcomer Outperforms Veterans
While Opus 4.8 ran with the most rules and analyses—over 80 learned rules—its discipline faltered, and it left deals on the table. The fact that Kimi K3, without any special effort parameter (API default), managed to outperform longstanding models suggests that model robustness and decision integrity matter more than raw complexity.
The Fairness Note
It’s important to mention that Kimi K3 ran without an effort parameter (the API default), while the other models operated at a high effort setting, known as xhigh. This makes K3’s performance even more impressive in a fair comparison.
As an affiliate, we earn on qualifying purchases.
The Big Takeaway for Business and AI
In the end, the question for companies considering AI tools isn’t just whether they can generate compelling chat responses. It’s whether they can finish what they start, resist manipulation, and read deeply into documents—traits that truly determine whether AI can be trusted to run real operations. The league table and experiment results at firmulate.com/benchmarks.html make this clear: choosing the right AI model is now a strategic decision, not a game of chance.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
