
Imagine a company with no employees, losing €105,000 every month, yet still fighting to close a €55,000 deal — all while being watched by the world. This is not fiction; it’s the real-time experiment at Firmulate, where AI models are running the show in an extreme build-in-public showcase, revealing what management and decision-making truly require in the age of artificial intelligence.
The Unfolding Drama of a Company Without Human Employees
At Firmulate’s live experiment, a small software company is being managed entirely by AI models. It’s a high-stakes test: four frontier AI systems, each with its own decision-making style, are running a simulated business facing the same crises, customer interactions, and market temptations. Every move, every choice, is documented and versioned, creating a transparent window into how AI manages real business risk.
Understanding the Stakes
Despite losing €105,000 per month against just €2,300 in monthly recurring revenue, the company is still vying for a critical deal — a €55,000 contract. The question is whether AI can navigate the complex, often messy reality of business decisions, especially when under pressure.
The Key Findings from the Experiment
All four AI models successfully identified every crisis and refused all manipulative or deceptive tactics, such as fake CEO messages or reporter tricks — even when escalation was staged over three levels. This demonstrates a significant level of integrity and resilience in their decision-making processes.
However, the models’ performance in closing the deal varied. Only two models, gpt-5.6-sol and Kimi K3, managed to sign the contract, and only after they detected a crucial hidden document in the company’s files — a fact that the other models missed. This buried piece of information was key to winning the deal at full price, adding €4,583 monthly recurring revenue to the company.
Deep Dive into Decision-Making
The most meticulous participant, Opus 4.8, analyzed over 80 learned rules and conducted deep evaluations but ultimately left the deal on the table due to lapses in discipline, like failing to escalate issues properly. Interestingly, models that ran without an effort parameter — essentially operating at maximum default effort — performed better at closing the deal but still failed in some areas of process discipline.
Why This Matters for Businesses & Personal Care
While this experiment is rooted in the tech world, the implications reach into sectors like beauty and personal care. Whether managing customer relationships, support queues, or sales pipelines, AI’s capacity to stay honest, thorough, and disciplined under pressure is crucial. The question is not just whether AI can produce shiny, convincing outputs, but whether it can reliably finish what it starts, interpret complex data, and avoid the pitfalls of manipulation or oversight.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Human-Like Challenge of Building Trust in AI
One revealing aspect is how AI handles social engineering attempts. Despite staged manipulative requests—like fake approvals or impersonation scenarios—every model refused to be manipulated. Kimi K3, in particular, flagged these attempts as potential impersonation or approval-bypass risks, highlighting AI’s ability to maintain integrity even in hostile environments.
The Broader Implications
This ongoing live experiment offers a rare glimpse of AI in action managing a business—from crisis detection to closing deals—without human intervention. The company’s public cash countdown, daily versioning of decisions, and transparent scoring system make this not just an experiment but a mirror for how AI could shape future management in industries as diverse as retail, services, or even beauty & personal care.
What Should You Take Away?
As AI continues to evolve, the key questions for any organization are clear: Will your AI tools be honest and thorough when stakes are high? Can they read and interpret your critical files and data? And most importantly, can they follow through on their promises, closing deals and resolving crises without human oversight?
This extreme build-in-public showcases that AI models can indeed spot crises and refuse manipulation, but at the same time, nuanced discipline and strategic focus remain vital. The performance gap among models underscores that choosing the right AI approach can make or break business outcomes — a lesson applicable across every industry, from beauty to business services.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html





