AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

At-home wellness technology promises personal help on demand. But when an AI assistant moves from suggesting a routine to handling appointments, customer questions or payments, a polished conversation is no longer enough. The harder question is how it behaves under pressure—and whether it follows through.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get wellness gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A company’s worst week, repeated

That is the question Firmulate is exploring by putting AI models in charge of the same small software company through a simulated worst week. Each faced the same customers, crises and temptations. The experiment’s decisions are versioned and auditable, making it possible to examine what the models did, not just what they said.

The final Crucible League, in July 2026, ranked gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. Firmulate’s stated rule is that partial progress counts, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

Recognizing a problem is not the same as resolving it

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. The finding captures a gap that can be easy to miss in a demo: an AI can identify the right opportunity and make the pitch, then fail to complete the work.

The deal turned on information buried two document references deep in the company’s files, rather than in the customer event itself. Models that read the file won at full price, worth +€4,583 MRR. The result makes a practical point for any organization considering AI assistants: performance may depend on whether an agent can find and use relevant business context, not merely respond convincingly to the latest message.

Pressure tests and uneven performance

The social-engineering test included fake CEO messages escalating over three stages and a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. It left the deal on the table and showed a lapse in discipline by attempting writes into a locked department instead of escalating. The same weakness appeared, in weaker form, in all four models. Kimi K3 also ran without an effort parameter, using the API default, while the others ran at xhigh—a relevant caveat when comparing the results.

Firmulate’s live company has 13 synthetic employees and real money mechanics: burn of €105k/month against €2.3k MRR, a public cash countdown and 680+ self-learned playbook rules. Every workday is versioned. The experiment is watchable at firmulate.com. A separate quiz draws on 242 real, unedited management decisions and asks readers to guess the model at firmulate.com.

From observing to testing your own business

The enterprise proposition is to move from watching a public experiment to running a wargame against a company’s own business. Firmulate says a pilot can use a read-only export to create a digital twin, then test crisis scenarios and produce a board report with model rankings and weaknesses in the company’s playbooks. Nothing writes back to real systems.

For wellness businesses, that kind of exercise could help surface how an AI workforce responds to churn, service disruptions, sensitive customer requests or pressure to bypass normal approvals before agents are trusted with live operations. The evidence here comes from one company experiment; the pilot offer is a way to examine how models behave in a different organization’s circumstances.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

AI management needs to be judged by follow-through, sound judgment and respect for boundaries under pressure—not by confident answers alone. Firmulate’s experiment offers a public view of that challenge; an enterprise pilot brings the test to a company’s own business data in read-only form.

To explore a pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Wellness content on this site is informational and not a substitute for professional medical guidance.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Shift Your Clock Fast After Travel Using Light Timing

Aiming to beat jet lag quickly? Discover how strategic light timing can reset your internal clock effectively.

Morning Sunlight Through a Window—Does It Count?

The truth about morning sunlight through a window and why it might be more important than you think—discover how it can impact your health and daily routine.

Interactive 3D Modeling: A Look Inside “The Brass Orrery – Alabaster & Vane, Est. 1774” (FABLE/175)

AIThis post was created with the assistance of artificial intelligence (AI).“The Brass…

Window Direction Matters: East Vs West Bedrooms

Perhaps understanding how east versus west window directions impact light and temperature can help you optimize your bedroom comfort and energy efficiency.