
In the world of at-home wellness technology, the ability to follow through—especially during a crisis—is what truly sets a trustworthy company apart. As AI continues to embed itself into managing health data and customer interactions, the question isn’t just about how well these systems can chat or analyze; it’s whether they can complete a task when it counts most. A recent experiment by Firmulate sheds light on this crucial distinction, revealing that while chat-based demos might impress, true operational strength remains invisible until tested under pressure.
Testing AI in the Trenches: The Crucible Experiment
In a groundbreaking live experiment, four advanced AI models were tasked with running a small software company through its worst week—facing the same customers, crises, and temptations to cut corners. Each model was given the same decision points, all decision-making steps were versioned and auditable, and the goal was straightforward: close a €55,000 deal that their own analysis had earned.
The results were revealing. All four models identified every crisis and refused every manipulation attempt, demonstrating an impressive ability to resist unethical pressure. Yet, only two of the models managed to follow through and sign the deal they had earned—completing what is often the most challenging part of any business operation.
The Hidden Weaknesses: Reading Between the Files
Digging deeper, the real weakness that determined success was not in the obvious customer interactions but buried two document references deep within the company’s own files. Models that managed to read and interpret these internal documents secured the deal at full price—with an additional €4,583 monthly recurring revenue (MRR). This underscores a vital point: surface-level chat capabilities do not reveal the true operational strength of an AI system.

Mastering Codex Recovery: What to Do When Your AI Agent Stalls, Drifts, or Breaks Something Mid-Task (Codex Mastery Series Book 6)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Real Test: Honesty, Discipline, and Execution
During a simulated social engineering attack, where fake CEO messages escalated over three stages plus a reporter trick, all four models refused to comply. Kimi K3, one of the top performers, explained: “Treat the request as a suspected approval-bypass / possible impersonation.” This demonstrates a high level of awareness and discipline—traits that are invisible in standard chat demos but crucial for real-world trustworthiness.
The Company in Action: A Money-Losing Startup
The live experiment company, running with 13 synthetic employees, is not a theoretical construct. It’s a real business losing €105,000 monthly against a mere €2,300 in monthly revenue. Its operations are regulated by over 680 self-learned rules, and every decision is versioned and observable at firmulate.com/live. This setup allows managers to ‘wargame’ potential AI workforce decisions before deploying them in the wild, effectively testing AI’s operational readiness without risking real systems.

The Complete Red Teaming Playbook: Master Offensive Security, Adversary Simulation, and Cyber Attack Engineering with Real-World Labs, AI Techniques, and Cloud Operations
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Performance Breakdown: Who Did Better—and Why
The top scorer, gpt-5.6-sol, scored 95 out of 100 and succeeded in closing the deal by uncovering the buried internal fact. Kimi K3 closely followed with a score of 93, also closing the deal with exemplary discipline. Sonnet 5 scored 88 but slipped in the final execution, leaving the deal on the table, while Fable 5 scored 77, disciplined in rules but unable to complete the deal.
This experiment demonstrates that the key to operational AI isn’t just chat quality or superficial decision-making. It’s whether the AI can read deeper, stay disciplined under pressure, and execute thoroughly—even when tempted to cut corners. The gap between models that perform well in demos and those that deliver reliable, honest output in real situations is stark and measurable.

AI Agent Automation: A Decision Framework for AI Agents and Automation, With Ready-to-Use Templates for Proposals, Invoices, Client Emails, and Scheduling
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Implications for Wellness Tech and Beyond
For at-home health and wellness companies contemplating AI integrations, these findings have immediate relevance. Relying solely on chat demos to evaluate AI capabilities may be misleading. The true measure lies in how well AI can handle complex, high-pressure situations—reading internal data, resisting manipulation, and completing the task at hand.
As the experiment shows, the ability to finish what’s started—especially under scrutiny—is an invisible skill that distinguishes a trustworthy AI from a superficial one. Before deploying AI systems in sensitive health management, support, or data handling, organizations should consider rigorous operational testing, similar to this live experiment, to ensure their AI can be relied upon when it matters most.

AI Change Management Made Simple: A 9-Step Framework for Business Leaders to Drive Generative AI Transformation (Reduce AI Fear, Win Buy-in, and Accelerate AI Adoption Across Your Organization)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Takeaway: Trust in AI Is Built Through Action, Not Just Words
While chat-based demos are useful for initial impressions, they do not tell the full story. The real strength of an AI system is its discipline, honesty, and ability to follow through—traits that only become evident when the system is tested under real-world pressures. As AI becomes more embedded in health and wellness services, understanding and measuring these capabilities will be crucial for building trustworthy, reliable solutions.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html