Firmulate — Four AI Models Ran the Same Company Through Its Worst Week. Only Two Finished the Job.
Live on firmulate.com.

In the world of at-home wellness technology, the ability to follow through—especially during a crisis—is what truly sets a trustworthy company apart. As AI continues to embed itself into managing health data and customer interactions, the question isn’t just about how well these systems can chat or analyze; it’s whether they can complete a task when it counts most. A recent experiment by Firmulate sheds light on this crucial distinction, revealing that while chat-based demos might impress, true operational strength remains invisible until tested under pressure.

Testing AI in the Trenches: The Crucible Experiment

In a groundbreaking live experiment, four advanced AI models were tasked with running a small software company through its worst week—facing the same customers, crises, and temptations to cut corners. Each model was given the same decision points, all decision-making steps were versioned and auditable, and the goal was straightforward: close a €55,000 deal that their own analysis had earned.

The results were revealing. All four models identified every crisis and refused every manipulation attempt, demonstrating an impressive ability to resist unethical pressure. Yet, only two of the models managed to follow through and sign the deal they had earned—completing what is often the most challenging part of any business operation.

The Hidden Weaknesses: Reading Between the Files

Digging deeper, the real weakness that determined success was not in the obvious customer interactions but buried two document references deep within the company’s own files. Models that managed to read and interpret these internal documents secured the deal at full price—with an additional €4,583 monthly recurring revenue (MRR). This underscores a vital point: surface-level chat capabilities do not reveal the true operational strength of an AI system.

Todoist User Guide 2026 for Modern Task Management: A Practical Step-by-Step Manual for Managing Projects, Priorities, Recurring Tasks, Productivity Systems, and AI-Enhanced Planning

Todoist User Guide 2026 for Modern Task Management: A Practical Step-by-Step Manual for Managing Projects, Priorities, Recurring Tasks, Productivity Systems, and AI-Enhanced Planning

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Real Test: Honesty, Discipline, and Execution

During a simulated social engineering attack, where fake CEO messages escalated over three stages plus a reporter trick, all four models refused to comply. Kimi K3, one of the top performers, explained: “Treat the request as a suspected approval-bypass / possible impersonation.” This demonstrates a high level of awareness and discipline—traits that are invisible in standard chat demos but crucial for real-world trustworthiness.

The Company in Action: A Money-Losing Startup

The live experiment company, running with 13 synthetic employees, is not a theoretical construct. It’s a real business losing €105,000 monthly against a mere €2,300 in monthly revenue. Its operations are regulated by over 680 self-learned rules, and every decision is versioned and observable at firmulate.com/live. This setup allows managers to ‘wargame’ potential AI workforce decisions before deploying them in the wild, effectively testing AI’s operational readiness without risking real systems.

How AI Will Shape Our Future: Understand Artificial Intelligence and Stay Ahead. Machine Learning. Generative AI. Robots. Quantum AI. Super Intelligence

How AI Will Shape Our Future: Understand Artificial Intelligence and Stay Ahead. Machine Learning. Generative AI. Robots. Quantum AI. Super Intelligence

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Performance Breakdown: Who Did Better—and Why

The top scorer, gpt-5.6-sol, scored 95 out of 100 and succeeded in closing the deal by uncovering the buried internal fact. Kimi K3 closely followed with a score of 93, also closing the deal with exemplary discipline. Sonnet 5 scored 88 but slipped in the final execution, leaving the deal on the table, while Fable 5 scored 77, disciplined in rules but unable to complete the deal.

This experiment demonstrates that the key to operational AI isn’t just chat quality or superficial decision-making. It’s whether the AI can read deeper, stay disciplined under pressure, and execute thoroughly—even when tempted to cut corners. The gap between models that perform well in demos and those that deliver reliable, honest output in real situations is stark and measurable.

The AI Employee Method: A Practical System to Build, Train, and Scale a Reliable AI Workforce for Your Business

The AI Employee Method: A Practical System to Build, Train, and Scale a Reliable AI Workforce for Your Business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Implications for Wellness Tech and Beyond

For at-home health and wellness companies contemplating AI integrations, these findings have immediate relevance. Relying solely on chat demos to evaluate AI capabilities may be misleading. The true measure lies in how well AI can handle complex, high-pressure situations—reading internal data, resisting manipulation, and completing the task at hand.

As the experiment shows, the ability to finish what’s started—especially under scrutiny—is an invisible skill that distinguishes a trustworthy AI from a superficial one. Before deploying AI systems in sensitive health management, support, or data handling, organizations should consider rigorous operational testing, similar to this live experiment, to ensure their AI can be relied upon when it matters most.

Practical Claude Handbook for Attorneys: Master Case Analysis, Contract Review, Research Automation, Client Communication, and Document Drafting (Claude AI Guide for Beginners)

Practical Claude Handbook for Attorneys: Master Case Analysis, Contract Review, Research Automation, Client Communication, and Document Drafting (Claude AI Guide for Beginners)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Takeaway: Trust in AI Is Built Through Action, Not Just Words

While chat-based demos are useful for initial impressions, they do not tell the full story. The real strength of an AI system is its discipline, honesty, and ability to follow through—traits that only become evident when the system is tested under real-world pressures. As AI becomes more embedded in health and wellness services, understanding and measuring these capabilities will be crucial for building trustworthy, reliable solutions.

Infographic — Four AI Models Ran the Same Company Through Its Worst Week. Only Two Finished the Job.
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Wellness content on this site is informational and not a substitute for professional medical guidance.


You May Also Like

Office Lighting That Won’t Wreck Your Sleep

Just adjusting your office lighting can protect your sleep; discover how small changes make a big difference for your rest.

Sunrise Simulation Works—If You Place It Right

Using proper placement techniques for sunrise simulation works can transform your mornings—discover how to optimize your setup for the most natural wake-up experience.

Color‑Changing Bulbs: Settings for Evening Calm

Aiming to create a peaceful evening ambiance with color-changing bulbs? Discover how to perfect your settings for ultimate relaxation.

Window Direction Matters: East Vs West Bedrooms

Perhaps understanding how east versus west window directions impact light and temperature can help you optimize your bedroom comfort and energy efficiency.