firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get baking supplies delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

What Ice Cream and AI Management Have in Common?

Just as choosing the right ingredients makes your signature ice cream stand out, selecting the best AI model can determine whether your business thrives or dives. Recent experiments show that behind the scenes of AI decision-making are crucial qualities that matter — honesty, discipline, and the ability to finish what it starts. And in a real-world test, the newcomer among AI models managed to beat some of the most established players, all without any special effort parameter.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Real-World AI Company Experiment

At Firmulate, a live and watchable experiment pits four leading AI models against one another—each tasked with running a small, real software company through its toughest week. This is no simple demo; it’s a fully operational test where every decision and action is tracked, versioned, and auditable. The goal? To see which AI can manage crises, resist manipulation, and ultimately close profitable deals, just like a human executive.

The League Table of Performance

  • gpt-5.6-sol scored 95, leading the pack by finding buried information in the company files and closing the deal.
  • The newcomer Kimi K3 scored 93, narrowly behind, demonstrating the cleanest discipline of the field and winning the same deal — at full price.
  • Sonnet 5 scored 88, also closing the deal but with a few process slips.
  • Fable 5 scored 77, while Opus 4.8 finished with 73, both closing the deal despite a few weaknesses.

Remarkably, all models identified every crisis and refused manipulative tactics, such as fake CEO messages or reporter tricks. Yet, only Kimi K3 and gpt-5.6-sol actually signed the €55,000 deal their own analysis justified—showing the importance of disciplined decision-making.

The Hidden Weakness in the Files

The decisive advantage for Kimi K3 and gpt-5.6-sol was their ability to read and analyze company files to uncover buried information that wasn’t obvious from the customer’s immediate story. This subtle but crucial insight made the difference between just closing a deal and closing it at full value—adding €4,583 in monthly recurring revenue.

Discipline Under Pressure

The experiment also tested resilience against sleight-of-hand tactics. All models refused to accept fake approvals or background requests to bypass protocol, with Kimi K3 explicitly reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This disciplined stance is vital in real business, where manipulative tactics are common.

The Real Business Is Live and Losing Money

The company managed by these models is not a simulation; it’s a live operation with 13 synthetic employees, burning €105,000 monthly against €2,300 in monthly revenue. Every workday, the decision-making processes are versioned and transparent, providing a unique window into how AI can run real companies. Watch this ongoing experiment at firmulate.com/live.

Amazon

business AI management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Surprising Result: A Newcomer Outperforms Veterans

While Opus 4.8 ran with the most rules and analyses—over 80 learned rules—its discipline faltered, and it left deals on the table. The fact that Kimi K3, without any special effort parameter (API default), managed to outperform longstanding models suggests that model robustness and decision integrity matter more than raw complexity.

The Fairness Note

It’s important to mention that Kimi K3 ran without an effort parameter (the API default), while the other models operated at a high effort setting, known as xhigh. This makes K3’s performance even more impressive in a fair comparison.

Amazon

AI crisis management platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Big Takeaway for Business and AI

In the end, the question for companies considering AI tools isn’t just whether they can generate compelling chat responses. It’s whether they can finish what they start, resist manipulation, and read deeply into documents—traits that truly determine whether AI can be trusted to run real operations. The league table and experiment results at firmulate.com/benchmarks.html make this clear: choosing the right AI model is now a strategic decision, not a game of chance.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI deal closing software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Our Favorite Savory Scented Candles, Including One That Smells Just Like Freshly Baked Bread

Explore our favorite savory-scented candles, including one that mimics the smell of freshly baked bread, now available for home fragrance enthusiasts.

National French Fry Day Surges In Global Coverage

Coverage of National French Fry Day has surged worldwide, with 25 mentions in recent media reports, highlighting increased popularity and marketing efforts.

Why Sweeteners Don’t Behave Like Easy Swaps

Genuine insights reveal why sweeteners often fall short as simple substitutes, leaving lingering questions about their true effects and safety.

Why Flour Age Changes Baking Performance

Baking quality declines as flour ages, affecting gluten strength and dough performance—discover how to adapt for perfect results every time.