
Imagine running your favorite ice cream shop, day after day, with every decision made by artificial intelligence. But instead of sweet success, this shop is bleeding €105,000 every month—yet you can see it live, every single day. Welcome to an unprecedented experiment where AI models are running a real software company in public, facing crises, temptations, and tough choices, all while their performance is scrutinized in real-time.
The Live Company That’s Changing What We Know About AI
At first glance, it looks like just another company, but this one is unlike any you’ve seen. It’s a real, functioning software business—burning through €105,000 each month against a modest €2,300 in monthly recurring revenue, with a public countdown clock showing how long it can keep going. Every day, it’s managed by 13 synthetic employees, guided by over 680 self-learned rules that adapt and evolve through every workday.
This company is part of a groundbreaking experiment by Firmulate, an initiative designed to test how AI models handle real-world management challenges, crisis responses, and ethical dilemmas—under the harsh spotlight of live performance. Each decision is meticulously recorded, versioned, and made auditable, offering a rare window into AI’s true capabilities and limitations in a complex, money-driven environment.
The AI Models and Their Performance
- Four frontier AI models are tested: gpt-5.6-sol, Kimi K3, Sonnet 5, and Opus 4.8.
- Scores range from 77 to 95, with the highest score achieved by gpt-5.6-sol.
- All models identified every crisis and refused manipulation attempts, demonstrating honesty and situational awareness.
One of the most revealing findings: the models that read and interpret internal company files closed the most lucrative deal, worth +€4,583 in monthly recurring revenue. The secret to this success was buried two document references deep in the company’s own files—an insight that was only uncovered by models that thoroughly examined the company’s internal knowledge base.
Ethical Resilience Under Pressure
Actors tried to deceive the AI—sending fake CEO messages escalating over multiple stages and even attempting a reporter trick with a backchannel yes/no query. Remarkably, all five AI models refused these manipulations, with Kimi K3 explicitly treating such requests as potential impersonation or approval-bypass attempts.

AI in Property Management: A Practical, Unboring Look at Artificial Intelligence in the Multifamily Industry
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Challenges of Running a Fake Business Live
This experiment isn’t just a tech showcase; it’s a brutal test of AI discipline and decision-making under pressure. The company, despite its AI-driven management, is losing money daily, illustrating how hard it is for even sophisticated algorithms to sustain real-world operations without human oversight.
The Opus 4.8 model, with the deepest analysis and over 80 learned rules, showed the most thorough processing but still faltered at the last mile—failing to close a deal due to discipline slips and a tendency to leave opportunities on the table rather than escalate. This highlights a crucial point: even the best AI models can struggle with nuanced management tasks when faced with incomplete information or internal conflicts.
What This Means for Business and AI
For anyone managing customer relationships, forecasting, or operational decision-making, the question is no longer just about AI’s ability to generate convincing chat responses. It’s about whether AI can truly finish what it starts, read and interpret internal company data, and stay honest when tempted to cut corners. As this experiment demonstrates, these qualities are essential for AI to be genuinely trustworthy in business environments.

Hands-On Simulation Modeling with Python: Develop simulation models for improved efficiency and precision in the decision-making process, 2nd Edition
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why You Should Follow This Live Experiment
Unlike typical AI demos, this is a real, running company—every weekday, making decisions, failing, succeeding, and losing money live. You can watch its progress, examine decisions, read actual internal messages, and even guess which model made which decision at firmulate.com/live.html. It’s transparency in its purest form, offering an unprecedented view into AI’s management skills in a high-stakes environment.
Key Takeaways for Business Leaders
- AI can identify and respond to crises effectively—if it’s designed and trained properly.
- Honesty and resistance to manipulation are achievable traits, even in complex scenarios.
- Understanding an AI’s internal knowledge—like buried internal data—can be the difference between winning or losing a critical deal.
- Despite all the progress, AI still struggles with discipline and closing opportunities—lessons that matter for real-world deployment.
For executives and technologists alike, this experiment underscores a vital truth: AI management tools must be tested rigorously, not just in chat demos but in the real chaos of business. The future of AI in enterprise lies in transparency, discipline, and the ability to complete what it starts—lessons that this live, losing company is painfully illustrating every day.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

As an affiliate, we earn on qualifying purchases.

People Analytics: Using data-driven HR and Gen AI as a business asset
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.