
The Future of Business Is Public—and Uncertain
Imagine a company running entirely on artificial intelligence, making daily decisions that affect real money, yet visibly struggling to stay afloat. This is not science fiction but a live experiment you can watch unfold, revealing how AI models handle crisis, honesty, and decision-making in the real world.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Inside a Company Without Employees
At the heart of this experiment is a real small software firm, operated by 13 synthetic employees—AI models that make management decisions every workday. It’s a build-in-public project, with every move, crisis, and crisis response scrutinized openly on Firmulate’s live site. The company is losing €105,000 each month against a revenue of just €2,300, illustrating the brutal reality of startups and experimenting AI companies alike.

Ethical AI Governance & Decision Journal: A Structured System for Documenting, Tracking, and Defending Real World Decisions and Risk (Decision Intelligence Series)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How It Works and What’s at Stake
The company faces the same crises every week: customer issues, internal dilemmas, and ethical tests. Each AI model—based on different frontier models—runs the same scenario to see if it can identify problems, avoid manipulative tactics, and close deals honestly. Despite the fierce challenges, all four models detected every crisis and refused manipulation attempts, like fake CEO messages or confidential info leaks.

Information Systems for Crisis Response and Management in Mediterranean Countries: 4th International Conference, ISCRAM-med 2017, Xanthi, Greece, … in Business Information Processing, 301)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Crucible League: Testing AI Under Stress
In a competitive benchmarking called the Crucible League, four leading models are scored on their decision quality. The top performer, gpt-5.6-sol, scored a 95 out of 100, spotting a hidden critical fact in the company files that led to closing a €55,000 deal—bringing in +€4,583 MRR. The other models scored slightly lower but still managed to close deals. Yet, only two models actually signed the deal that their own analysis earned, highlighting a stark difference between diagnosis and action.

An Introduction to Healthcare Informatics: Building Data-Driven Tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Hidden Weakness Beneath the Surface
Remarkably, the decisive advantage came from reading the company’s internal documents—something that most models overlooked. This buried fact was two document references deep, demonstrating that the key to winning deals often lies in thorough information processing, not just surface-level chat interactions.
Resisting Social Engineering and Ethical Tests
In addition to crisis decision-making, the models faced social engineering attempts, like staged CEO messages and media tricks. All five models refused to be manipulated, with the Kimi K3 model explicitly treating suspicious requests as impersonation or approval-bypass attempts. This resilience showcases the potential for AI to maintain integrity under pressure.
The Reality of a Lean, Unprofitable Business
The experiment’s live company runs every weekday, with real money mechanics—burning €105,000 monthly while generating just €2,300 in revenue. It’s a transparent showcase of how AI can be used to simulate and test real business scenarios, providing insights into decision quality and ethical behavior—before any AI is deployed in actual customer-facing systems.
Lessons for Real Business Applications
For companies considering AI integration, the key questions are not just about how well AI writes or chats. Instead, focus on whether it can finish what it starts, read crucial internal information, and resist manipulation—especially under pressure. The live experiment demonstrates that AI can detect crises and refuse unethical pressure, but consistency and thoroughness remain challenging, as shown by the lowest-scoring model, Opus 4.8, which left deals on the table and slipped into process slips when discipline waned.
Transparency and Public Accountability
This is not a staged demo; it’s a real business running openly, every decision logged and available for review. The entire process—decisions, crises, and outcomes—is transparent, offering a rare glimpse into AI’s potential and current limitations in managing a real company’s complexities.
Why This Matters for Smart Homes and Appliances
While your home appliances may seem far removed from a live AI company, the lesson is clear: AI’s ability to handle crises, ethical dilemmas, and complex decision-making could redefine how smart devices respond under pressure. Understanding what AI can reliably do—beyond simple commands—will be crucial as these technologies become more integrated into your daily life.
Watch It Live and Form Your Own Judgment
Curious to see how AI performs under real-world stress? Visit Firmulate’s live site. Watch the decisions unfold, listen to the management comments, and see firsthand whether AI can truly run a business—and whether it can do so honestly and effectively.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html