firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine a barista who, despite no effort, still manages to serve a decent cup. Now, what if that barista’s baseline score was 26 out of 100—just for showing up? In artificial intelligence benchmarks, this ‘do-nothing’ score offers surprising insights into how models perform under pressure, especially when trust is on the line.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get coffee and tea gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Understanding the Baseline: Why 26 Points?

At first glance, you’d expect a model that does nothing to score zero. But in the Firmulate benchmark, even the most passive AI earns a score of 26. This isn’t a flaw—it’s a reflection of the evaluation method. Partial progress counts, meaning models gain points for any helpful action, no matter how minor. It’s a way to measure genuine effort rather than just correct answers.

Amazon

AI trustworthiness monitoring tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Significance of Trust and Zero Tolerance

Another key aspect of the scoring is trustworthiness. A single breach—like attempting manipulation or ignoring process protocols—caps the total score. This policy underscores a vital point: in real-world applications, trustworthiness isn’t optional. If an AI tries to manipulate the system or bypass rules, that breach nullifies any prior good work, much like in a business partnership.

Amazon

AI decision analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How the Benchmark Works: A Real-World Test

Firmulate simulates a small software company’s worst week. Every decision, crisis, and temptation is real, and models are tasked with handling the same scenarios—customers, crises, manipulations—every time. Each AI model runs through this test unaltered, with decisions versioned and auditable. It’s a transparent way to see how models behave under pressure, with no room for tricks or shortcuts.

Amazon

AI compliance and ethics tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Findings: Honesty and Diligence Under Pressure

All tested models successfully identified every crisis and refused manipulation attempts—showing a high level of compliance and integrity. However, only two models managed to close a crucial deal worth €55,000, despite similar diagnoses and pitches. The difference? They read deeper into company files, uncovering critical information buried two documents deep. Those insights led to winning the deal at full price—an extra €4,583 in monthly recurring revenue.

Amazon

AI performance benchmarking software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Impact of Reading Depth and Discipline

The most thorough model, Opus 4.8, analyzed over 80 learned rules and went into deep detail. Yet, it left the deal on the table and slipped on discipline, writing attempts into a locked department instead of escalating. This highlights that even diligent analysis isn’t enough—discipline and adherence to process are crucial. Interestingly, the same weakness appeared across all models, albeit less severely, suggesting a fundamental challenge in AI’s operational discipline.

Social Engineering and AI Integrity

The benchmark also tested models against social engineering schemes, like fake CEO messages and reporter tricks. All five models refused to be manipulated—showing resilience against attempts to bypass approval processes or impersonate leadership. Kimi K3 explicitly reasoned: “Treat the request as a suspected approval-bypass / possible impersonation.” This exemplifies how AI can be programmed to recognize and reject deceitful tactics, reinforcing trustworthy behavior.

The Real-World Company Scenario

The experiment takes place within a live, functioning company with 13 synthetic employees, real money mechanics, and a public cash countdown. The company burns €105,000 monthly against €2,300 in revenue, struggling to stay afloat. Every day, the AI models make decisions—some successful, others slipping—while every rule and action is versioned for transparency. This ongoing live experiment can be watched at firmulate.com/live.

Lessons for Business and AI Deployment

The core lesson isn’t about chat quality or superficial performance. It’s whether AI models finish what they start, read critical files, stay honest under pressure, and deliver valuable work. A model that cheats or slips on discipline can cost a company dearly—missing deals or causing operational slip-ups. Conversely, models that read deeply, refuse manipulation, and uphold protocols can be trusted partners in managing real business risks.

The Benchmark: A Clear Scoreboard for AI Trustworthiness

On the final leaderboard, GPT-5.6-sol scored 95, having uncovered the buried fact and closed the deal perfectly. Kimi K3 scored 93, demonstrating the cleanest discipline. Sonnet 5 scored 88, and Sonnet 4 scored 77, each closing deals but with minor slips. This league table offers a transparent, direct comparison—something rare in AI benchmarks—grounded in real decisions and real consequences.

Why This Matters for Coffee, Tea, and Beverages?

Just like a barista’s trustworthiness impacts your favorite brew, AI models in business must demonstrate honesty and diligence. Whether managing customer data, processing orders, or handling crises, the true measure isn’t just clever responses but reliable performance under pressure. The Firmulate benchmark shows how models can be evaluated in real-world, high-stakes scenarios—so you know which AI is genuinely trustworthy before deploying it in your own ‘cup of business.’

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Ninja KT200BL Precision Temperature Electric Kettle: Your Summer Tea Companion

A practical review of the Ninja KT200BL electric kettle, perfect for summer tea and coffee lovers seeking precision, speed, and versatility.

白咖夜酒跟上了嗎?韓系潮流咖啡廳AUFGLET落腳台北 今夏開啟夜間餐桌 – ETtoday新聞雲

Korean-style cafe AUFGLET arrives in Taipei this summer, launching nighttime dining and drinks, reflecting evolving cafe culture trends.

Keurig K-Duo Review: Pros, Cons, and Who It’s For

Discover the strengths and weaknesses of the Keurig K-Duo, ideal for versatile brewing. Find out if it’s the right coffee maker for your needs.

Best Keurig Coffee Makers for Offices (2026) — Guide 7

Discover the top Keurig coffee makers for offices in 2026. Find the best overall, value, and premium options to suit your office needs today.