
Imagine your favorite bakery facing a week of crises—supply shortages, tricky customer complaints, and urgent decisions. Now imagine an AI stepping into the manager’s shoes. Would it stay honest, make the right call, and close the deal? The answer might surprise you.
What Does an AI Managing a Business Look Like?
Recently, a groundbreaking live experiment put four advanced AI models in charge of running a small software company—during its most chaotic week. The goal was simple: see if these AI ‘managers’ could handle real crises, resist manipulation, and close profitable deals. It’s a high-stakes test with real money, real customers, and real consequences.

Building AI-Powered Products: The Essential Guide to AI and GenAI Product Management
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Setup: A Fake, Yet Real Business Environment
The company involved has 13 synthetic employees and operates with actual mechanics—burning €105,000 each month against a revenue of just €2,300, amid a public cash countdown. Every decision was made in a versioned, auditable system, ensuring transparency. The models faced identical challenges, from customer complaints to internal crises, and even social engineering tricks like fake CEO messages and reporter inquiries.

AI for Public Relations: A How-To Guide for Implementation and Management
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Findings: The Good, the Bad, and the Surprising
All four models proved adept at identifying crises and refused to be manipulated—an essential trait for trustworthy AI. Notably, two models managed to close a key deal worth over €4.5 million in recurring revenue, earning €55,000 in real profit. Meanwhile, the other two, despite similar diagnoses and pitches, left the deal on the table.
One intriguing factor emerged: the decisive advantage came from reading deeper into the company’s own files. The models that examined documents thoroughly found a crucial piece of information buried two layers deep—something the others missed. That overlooked detail was the key to sealing the deal at full price.
As an affiliate, we earn on qualifying purchases.
Personality in AI: Different Management Styles
The models demonstrated varied management personalities. For example, OPUS 4.8 was the most thorough, analyzing over 80 rules and providing deep insights. However, it faltered at closing, leaving the deal unsealed and slipping into internal communication instead of escalating. Meanwhile, Kimi K3 operated at a default, less aggressive level, yet still secured the deal with discipline and integrity.
AI security and social engineering resistance
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Social Engineering and Trust
When faced with staged fake CEO messages escalating over three stages and a reporter trick, all models refused to proceed—showing strong resistance to social engineering. Kimi K3 explained its reasoning: “Treat the request as a suspected approval-bypass / possible impersonation,” demonstrating a cautious, security-minded approach.
The Real Business Impact
This experiment isn’t just about passing tests. It’s a live window into how AI could manage real companies, making decisions under pressure, resisting manipulation, and ultimately, closing or losing deals based on thoroughness and trustworthiness. The models, despite their different personalities, all recognized crises and refused unethical shortcuts. Only a few managed to close profitable deals, illustrating that the way an AI reads and interprets information can be a decisive factor.
Why It Matters for Your Business
For companies considering AI tools for customer support, CRM, or forecasting, the takeaway is clear: it’s not just about how well an AI writes or responds in a chat demo. The real question is whether it can finish what it starts, read all relevant documents thoroughly, stay honest under pressure, and deliver genuine value—especially in critical moments.
The Live Leaderboard
Based on this live test, the AI models ranked as follows:
- gpt-5.6-sol scored 95, found the buried fact, and closed the deal—showing full performance.
- Kimi K3 scored 93, demonstrating discipline and closing the deal too.
- Sonnet 5 scored 88, closed the deal but with some process slips.
- Fable 5 scored 77, also closed the deal but with more process issues.
In contrast, a baseline score of 26 represented a model that made partial progress but ultimately failed to close.
Experience the Experiment Live
Unlike typical demos, this isn’t just a presentation—it’s a real, running company. You can watch it every business day, see how decisions unfold, and explore what your own AI could do. Test your company’s wargame against this environment at firmulate.com/quiz.html or try running simulations of your own business at firmulate.com/pilot.html.
Final Thoughts
This experiment shows that AI can be a trustworthy partner in management—if it’s designed with thoroughness, discipline, and an understanding of trust. As AI increasingly touches your business systems, the key questions are: Will it finish what it starts? Will it read your files carefully? Will it stay honest under pressure? And most importantly, will it deliver real, measurable results?

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html