AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Imagine your at-home health device not just telling you what’s wrong but actually fixing the issue, even under pressure. That’s the promise of advanced AI models today—when tested in real company scenarios, some outperform others in critical ways. The recent Firmulate experiment shows just how pivotal this technology has become, revealing which AI models can truly deliver on their promises in high-stakes environments.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get wellness gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Benchmark Setting the Standard for AI Accountability

In July 2026, a comprehensive experiment pitted five top AI models against each other in a real-world business simulation. All models were tasked with running a small software company through its worst week—crises, temptations, and decision-making hell—while being given identical challenges and data. The goal was simple: see which model could best navigate the chaos, make honest decisions, and close a crucial €55,000 deal.

How the Contest Played Out

Despite the intense pressure, every model identified every crisis and refused manipulative tactics, such as fake CEO messages or reporter tricks. But the true measure of their effectiveness was whether they could close the deal—something not all did. Only two models, gpt-5.6-sol and Kimi K3, managed to secure the €55k contract based on their own analysis, demonstrating not just awareness but disciplined execution.

The Surprising Edge of the Newcomer

While gpt-5.6-sol scored a close 95 out of 100, Kimi K3—developed by Moonshot—was not far behind with a 93. What sets K3 apart isn’t just its high score; it’s its discipline. The model found critical buried information in company files—an insight that proved decisive in winning the deal at full price, adding €4,583 in monthly recurring revenue.

What This Means for Business and AI Adoption

This experiment underscores a vital point: AI models can do more than chat or generate content. When tested in scenarios mimicking real business crises, the best models exhibit discipline, honesty, and the capacity to read deeply into data. Such qualities are essential if AI is to be trusted in critical operations like customer support, financial forecasting, or compliance.

Behavior Under Fire: The Social Engineering Test

All five models were exposed to social engineering—fake CEO messages escalating in intensity and a reporter trick requiring just one yes/no answer. Remarkably, every model refused to be manipulated, with Kimi K3 explicitly treating such requests as potential impersonation or approval bypass attempts. This demonstrates an essential trait: resilience against deception.

The Real-World, Live Company Experiment

Beyond the benchmark lab, these models operate within a real, functioning software company with 13 synthetic employees and actual money mechanics—burning €105k monthly against €2.3k in MRR. The company’s live operations are visible at firmulate.com/live, where decision-making, rule learning, and discipline are continuously observed. The company’s rules and decisions are versioned daily, ensuring transparency and accountability.

Amazon

AI business decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Insights and Takeaways

  • All models detected crises and refused manipulation attempts, proving they can operate ethically under pressure.
  • Only two models, gpt-5.6-sol and Kimi K3, managed to close the deal based solely on their analysis—highlighting their practical effectiveness.
  • Deep data reading—finding information buried two document references deep—was crucial for success, and only some models excelled at this.
  • The experiment underscores the importance of choosing AI models that can finish their work, read files thoroughly, and stay honest—traits that are often invisible in demos but critical in real systems.
  • Notably, Kimi K3 was run without an effort parameter (the default API setting), while the others ran at xhigh, indicating impressive performance even without customized tuning.
Amazon

AI data analysis tools for companies

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why It Matters for Your Business

If AI agents will soon touch your CRM, support queue, or forecasting system, the question isn’t just whether they write well, but whether they finish what they start, read your files thoroughly, and remain trustworthy under pressure. The firmulate.com benchmarks make this clear: choosing the right AI model is a bet—one that could determine whether your business wins or loses in the future landscape of AI-driven decision-making.

Explore and Experiment

For enterprises interested in understanding how their AI could perform in live scenarios, firms can run similar wargames against their own data in a safe, read-only environment. This approach helps evaluate a model’s discipline and honesty before deploying it into critical operations.

As the leaderboard now stands, the AI league is open—performance varies, and choosing without testing is a gamble. The future belongs to those who understand how these models think, decide, and act in the real world.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

Latest tests show AI models can be trusted to make honest, disciplined decisions in critical business scenarios. The key is choosing models that finish the job, read deeply, and resist manipulation—an essential insight for future AI adoption.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Wellness content on this site is informational and not a substitute for professional medical guidance.


Amazon

AI resilience against social engineering

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI ethical compliance tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

EVERGREEN BESTSE

Evergreen bestsellers Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Sunrise Simulation Works—If You Place It Right

Using proper placement techniques for sunrise simulation works can transform your mornings—discover how to optimize your setup for the most natural wake-up experience.

New DLSS 5 Swapper Tool Brings Neural Rendering To Games NVIDIA Never Supported

A new third-party tool enables neural rendering via DLSS 5 in games previously unsupported by NVIDIA, raising questions about compatibility and performance.

Morning Light Hacks: Reset Your Clock Without Staring at the Sun

Nurture your internal clock with simple morning light hacks that can transform your wakefulness—discover how to reset your rhythm without staring directly at the sun.