
Imagine running your bakery’s order system, but the AI only does the bare minimum, never one step beyond. Surprisingly, even a ‘do-nothing’ AI scores 26 out of 100 in a rigorous benchmark. That’s a wake-up call for anyone relying on automation — and it reveals how careful we must be when trusting AI with our businesses.
Get baking supplies delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Understanding the Baseline: Why Even Doing Nothing Gets You Points
In recent experiments conducted by Firmulate, a robust AI benchmarking platform, even a model that essentially does nothing but respond passively scores around 26 points. This score isn’t zero, because the benchmark measures more than just active decision-making; it accounts for partial progress and cautious responses. Think of it as a safety net—if the AI doesn’t attempt to manipulate or mislead, it earns some trust points, even if it doesn’t fully act.
AI decision-making software for small business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why Partial Progress Matters
In real-world scenarios—like managing a bakery’s supply chain or customer service—doing a little is better than doing nothing. The benchmark recognizes this by awarding partial points for responses that show basic awareness, even if they don’t resolve every crisis or opportunity. Conversely, a single breach of trust—such as attempting manipulation—caps the total score. This cap prevents overestimating an AI’s reliability based solely on good performance elsewhere.
AI ethics and trustworthiness tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Real Experiment: Running a Small Business Crisis Test
Firmulate’s experiment involved running four leading AI models through the same simulated week of a small software company facing real crises: customer complaints, financial deadlines, and ethical dilemmas. Each model was tasked with crucial decisions, from diagnosing issues to closing deals. All models identified every crisis and refused manipulative attempts, yet only two managed to close the deal and sign the €55,000 contract based on their own analyses.
AI security and social engineering prevention
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why the Results Matter for Business
This isn’t just about AI scoring high or low; it’s about trustworthiness and effectiveness. The models that refused manipulation showed integrity, but only those that read deeper into company documents and avoided shortcuts secured the complete deal. Surprisingly, the key weakness was not in customer interactions but in unexamined internal files—reading these files allowed the winning models to identify opportunities others missed, securing an additional €4,583 MRR.
As an affiliate, we earn on qualifying purchases.
The Human-Like Risk of Social Engineering
Another critical test involved social engineering attacks—fake CEO messages and reporter tricks. All models successfully refused these, showing a shared capacity for ethical resistance. Kimi K3’s developers noted their model treated these as impersonation or approval-bypass threats, which aligns with a cautious, security-first approach.
The Live Business Environment: An Ongoing Testbed
Firmulate’s live setup simulates a real business with 13 synthetic employees working with real finances—burning €105k a month against €2.3k monthly revenue, with a public cash countdown. Every day, the models make decisions based on over 680 learned rules, making the process transparent for observers. You can see this ongoing experiment at firmulate.com/live.
The Surprising Place of Opus 4.8 and What It Reveals
Among the models tested, Opus 4.8 was the most thorough, with over 80 learned rules and deep analyses, yet it placed last. It left deals on the table, hesitated, and slipped into departmental silos instead of escalating issues. This shows that more rules and analysis don’t necessarily translate into better performance if discipline and decision-making processes falter under pressure.
What This Means for Business Leaders
When choosing AI for your operations, the focus should not solely be on impressive chat responses or superficial capabilities. Instead, ask: Will it complete tasks reliably? Will it read critical documents? Will it stay honest in high-pressure situations? Benchmarks like this one from Firmulate provide a sober, transparent view of what AI can and cannot do—and why trustworthiness is paramount.
Final Takeaway: Trust Is a Cap, Not a Bonus
The experiment clearly shows that partial progress and integrity are valued more than merely making the right decision once. A single breach of trust caps the total score, emphasizing that honesty under pressure is non-negotiable. In practical terms, this means businesses should focus on AI models that demonstrate consistent honesty and thoroughness, not just quick wins or high scores in demos.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
