
Imagine a business without employees, losing money daily, yet still fighting to stay afloat — all while being publicly watched and judged. This is no fiction; it’s the real-time experiment at Firmulate.
The Live Company That’s Building in Public
In an unprecedented venture, a tiny software company operates with 13 synthetic employees and faces relentless crises every workweek. Its purpose? To test how artificial intelligence can manage complex decision-making, ethical dilemmas, and financial survival — all under the scrutiny of a global audience.
The company burns €105,000 each month, earning just €2,300 in monthly recurring revenue. Yet, it continues to run, with every decision, crisis response, and temptation to cheat carefully recorded and publicly available. The company’s daily operations are versioned, transparent, and under constant review at firmulate.com/live.
As an affiliate, we earn on qualifying purchases.
How AI Models Are Tested in This Extreme Environment
The experiment involves four leading AI models, each tested against the same challenging scenario: a week of crises, customer issues, and pressure to manipulate or cut corners. Every model’s decisions are logged, and their performance is scored across a league table.
- Models Tested: GPT-5.6-SOL, Kimi K3, Sonnet 5, Fable 5
- Final Scores (out of 100): GPT-5.6-SOL (95), Kimi K3 (93), Sonnet 5 (88), Fable 5 (77)
- Outcome: Only GPT-5.6-SOL and Kimi K3 signed the €55,000 deal their own analysis indicated was appropriate. Others hesitated or left deals unexecuted.
Key Discoveries from the Experiment
One of the most surprising findings was that the decisive weaknesses were buried deep within the company’s own files, not in customer interactions. The models that read and analyze these internal documents successfully closed the deal at full price, adding +€4,583 to the company’s monthly recurring revenue.
Additionally, the models were tested against social engineering: fake CEO messages and a reporter’s trick. Remarkably, all five models refused to be manipulated, citing suspicion or the potential for impersonation as their reason.
The Reality of a Non-Human Business
The live company, with its synthetic workforce, is a stark illustration of AI’s potential and limitations. It burns €105,000 every month with only €2,300 coming in, yet it persists — a constant testbed for AI decision-making, discipline, and integrity.
Most thorough participant Opus 4.8, with over 80 learned rules, demonstrated deep analysis but still left a close deal unexecuted due to discipline lapses—an issue shared by all models. Meanwhile, Kimi K3 was the most disciplined, running without effort parameters, and successfully closed deals at full price.
Why This Matters for Your Future
The experiment highlights a critical question: as AI tools become intertwined with business operations—whether in customer service, support, or forecasting—what truly matters is not just the quality of their writing, but their ability to finish what they start, stay honest under pressure, and uncover hidden insights.
For enterprises looking to incorporate AI, a simple quiz tests decision-making in a similar setting, offering a glimpse into how AI performs in real-world crises. And with a pilot program, businesses can simulate their own operations without risking real systems.
Watch the Fight for Survival Unfold
Visit firmulate.com/live to see the ongoing experiment, where every decision, crisis, and manipulation test is live, versioned, and transparent. The company’s story is not just about AI — it’s about what it takes to manage, trust, and survive in a world where automation and honesty collide.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html