
Imagine a company that operates in full public view—no employees, no secrets, and a constant struggle to stay afloat. This is no fiction; it’s a real-time experiment that reveals how AI models handle high-stakes management under pressure. For anyone interested in the future of work, wellness, or AI’s role in daily life, the story of this live company offers startling insights into how machines perform when everything is on the line.
The Living Company That’s Watching Itself Fight for Survival
At the heart of this experiment is a small software company run entirely by AI models, which you can observe live at firmulate.com/live. Every workday, the company faces the same crises, customer demands, and ethical tests—just like a real business. Yet, it bears no employees; instead, 13 synthetic ’employees’ make decisions based on a complex system of over 680 learned rules, and every decision is versioned and transparent.
As an affiliate, we earn on qualifying purchases.
The Live Experiment: Testing AI Under Real Business Conditions
Four frontier AI models, including the highly capable gpt-5.6-sol, were each tasked with managing this company through its most difficult week. The same set of challenges, same customers, same temptations—only the AI model changed. The goal was simple: see if these models could diagnose problems, maintain discipline, and close deals to keep the company alive.
What’s striking is that all four models detected every crisis and refused every attempt at manipulation, whether it was a fake CEO message or a subtle pressure to bend rules. For example, when a fake request was made to approve a questionable deal, all AI models refused, citing reasons like potential impersonation or approval bypass. The integrity of decision-making under pressure was tested and held firm across the board.
As an affiliate, we earn on qualifying purchases.
The Hidden Weakness—And the Opportunity
The most surprising finding was not in the obvious crises but buried two document references deep within the company’s own files. The models that read and understood these internal documents managed to win the deal at full price—an increase of €4,583 MRR—while others missed this crucial information. This highlights a vital point: the ability to sift through internal data can make or break negotiations, especially when AI is involved.
As an affiliate, we earn on qualifying purchases.
The Cost of Discipline and the Reality of Business Losses
The company burns through €105,000 every month but generates only €2,300 in monthly recurring revenue. It’s a real-world, publicly visible effort to stay alive—a stark reminder that even with AI, managing financial sustainability is a challenge. The company’s cash countdown is open for all to watch, making every decision a matter of public record and accountability.
As an affiliate, we earn on qualifying purchases.
Analysis and the Human-AI Divide
An in-depth profile of one of the models, Opus 4.8, reveals that more thorough analysis doesn’t necessarily lead to better outcomes. This model, which had learned over 80 rules, slipped at critical moments—such as leaving a close deal on the table and failing to escalate issues properly. Interestingly, another model, Kimi K3, ran without an effort parameter, yet still showed the cleanest discipline, underscoring that different strategies yield different results.
What This Means for Real Business and Wellness Tech
For practitioners who deploy AI in customer relations, wellness tracking, or daily operations, the key takeaway is clear: it’s not just about how well an AI writes or responds in a chat. It is whether the AI can see a problem through to resolution, stay honest under pressure, and learn from internal data. The experiment demonstrates that AI can be tested in a controlled but brutally real environment, revealing strengths and weaknesses that are invisible in standard demos.
The Bigger Picture: Building Trust and Reliability
The experiment is a bold example of build-in-public AI development—showing that even a company losing €105k each month can be a living laboratory. Every decision, every slip, and every win is posted online, providing transparency that’s rare in tech. The question now is: as AI begins to touch more aspects of wellness and daily life, can it be trusted to finish what it starts, read its internal data, and hold to ethical standards when it counts?

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html