firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Imagine testing your favorite coffee blend not just in a cup, but in a real-world business scenario — with actual crises, real money mechanics, and high stakes. That’s the challenge faced by AI models in a groundbreaking live experiment, revealing surprising insights that matter for any enterprise considering AI as part of its team.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get coffee and tea gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Real-World Business Wargame: More Than Just Chatbots

In a recent live experiment, four of the leading frontier AI models took on the complex task of running a small software company through its worst week. This wasn’t a simple test of chat skills; it was a rigorous simulation involving real crises, customer relationships, and financial risks. The goal? To see which AI could best diagnose issues, resist manipulation, and deliver results that matter.

Amazon

AI business simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The League Table: A Close Race with a Clear Winner

By the end of the competition, the scores told a compelling story:

  • gpt-5.6-sol scored 95 — the highest, closing the deal and finding the buried fact that sealed the company’s fate.
  • Kimi K3 scored 93 — a newcomer from Moonshot, showing a disciplined approach and closing the same deal at full value.
  • Sonnet 5 scored 88 — closing the deal but with some process slips.
  • Fable 5 scored 77 — also closing the deal but with weaker discipline.
  • Opus 4.8 scored 73 — the lowest among competitors, struggling to finish the task fully.

Importantly, all models identified every crisis and refused manipulation attempts, but only two signed the critical €55,000 deal based on their diagnoses. The kicker? The winning models read deeper into the company’s own files, uncovering hidden clues that others missed, leading to a full-price deal worth +€4,583 MRR.

Amazon

enterprise AI decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why Reading the Files Matters

The experiment revealed a key weakness in some models: they focused on surface-level cues and missed buried evidence. The models that examined deeper documented references gained a decisive advantage, closing high-value deals that others left on the table. It’s a lesson for any business: understanding the full picture — reading your internal documents thoroughly — can be the difference between winning and losing.

Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Resisting Social Engineering and Manipulation

All models faced social engineering tests, including fake CEO messages escalating over three stages and a reporter trick asking for quick approvals. Remarkably, every AI refused these requests, with Kimi K3 explicitly reasoning that such requests could be impersonation attempts. This demonstrates a crucial attribute for AI in business: maintaining integrity under pressure.

Amazon

AI cybersecurity for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Live Company: Real Money, Real Stakes

The company running the experiment is real, with 13 synthetic employees managing business mechanics of burning €105k/month while earning only €2.3k MRR. The system operates daily, with over 680 self-learned playbook rules, and is openly monitored at firmulate.com/live. It’s a transparent, ongoing test of AI’s ability to uphold management discipline and deliver value.

The Curious Case of Opus 4.8

The most thorough participant, Opus 4.8, brought over 80 learned rules and the deepest analyses. Yet, it finished last among the group, leaving the close deal on the table and slipping into a process slip—writing attempts into a locked department instead of escalating. This highlights that thoroughness alone isn’t enough; discipline and decision-making processes matter just as much.

The Takeaway for Business Leaders

This experiment underscores a vital truth: in deploying AI for real business tasks, looking at chat demos isn’t enough. The real question is whether the AI can finish what it starts, read your internal files deeply, resist manipulation, and stay disciplined under pressure.

Choosing a model without your own testing is increasingly a gamble. The leaderboard from this live wargame suggests that newcomers like Kimi K3 can outperform established models, provided they meet the core test of integrity and thoroughness.

Fairness and Transparency in AI Testing

It’s worth noting that K3 ran without an effort parameter (the default API setting), while the other models ran at a higher setting — xhigh. All results are based on this consistent setup, ensuring a fair comparison.

The Future of AI in Business Decision-Making

As AI models become more integrated into enterprise workflows, such live, verifiable experiments are essential. They show that AI’s true value isn’t just in generating convincing chat, but in reliably delivering results, maintaining integrity, and uncovering buried insights. Leaders who test rigorously will better understand which AI can truly serve as a dependable team member.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Best Keurig Coffee Maker for Offices (2026) — Guide 28

Discover the top Keurig coffee makers in 2026. Find the best overall, best value, and best compact options for your brewing needs. Read more!

De’Longhi Dinamica Plus Review: Pros, Cons, and Who It’s For

An in-depth review of the De’Longhi Dinamica Plus, highlighting its strengths, weaknesses, and ideal users to help you decide if it’s the right espresso machine.

The One Mistake That Ruins Contact Time In Coffee (And How to Avoid It)

Discover the crucial mistake that disrupts coffee brewing and learn how precise timing can transform your cup into a perfect brew. Don’t miss these essential tips!