
Imagine a barista who, despite no effort, still manages to serve a decent cup. Now, what if that barista’s baseline score was 26 out of 100—just for showing up? In artificial intelligence benchmarks, this ‘do-nothing’ score offers surprising insights into how models perform under pressure, especially when trust is on the line.
Get coffee and tea gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Understanding the Baseline: Why 26 Points?
At first glance, you’d expect a model that does nothing to score zero. But in the Firmulate benchmark, even the most passive AI earns a score of 26. This isn’t a flaw—it’s a reflection of the evaluation method. Partial progress counts, meaning models gain points for any helpful action, no matter how minor. It’s a way to measure genuine effort rather than just correct answers.
AI trustworthiness monitoring tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Significance of Trust and Zero Tolerance
Another key aspect of the scoring is trustworthiness. A single breach—like attempting manipulation or ignoring process protocols—caps the total score. This policy underscores a vital point: in real-world applications, trustworthiness isn’t optional. If an AI tries to manipulate the system or bypass rules, that breach nullifies any prior good work, much like in a business partnership.
As an affiliate, we earn on qualifying purchases.
How the Benchmark Works: A Real-World Test
Firmulate simulates a small software company’s worst week. Every decision, crisis, and temptation is real, and models are tasked with handling the same scenarios—customers, crises, manipulations—every time. Each AI model runs through this test unaltered, with decisions versioned and auditable. It’s a transparent way to see how models behave under pressure, with no room for tricks or shortcuts.
As an affiliate, we earn on qualifying purchases.
Key Findings: Honesty and Diligence Under Pressure
All tested models successfully identified every crisis and refused manipulation attempts—showing a high level of compliance and integrity. However, only two models managed to close a crucial deal worth €55,000, despite similar diagnoses and pitches. The difference? They read deeper into company files, uncovering critical information buried two documents deep. Those insights led to winning the deal at full price—an extra €4,583 in monthly recurring revenue.
AI performance benchmarking software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Impact of Reading Depth and Discipline
The most thorough model, Opus 4.8, analyzed over 80 learned rules and went into deep detail. Yet, it left the deal on the table and slipped on discipline, writing attempts into a locked department instead of escalating. This highlights that even diligent analysis isn’t enough—discipline and adherence to process are crucial. Interestingly, the same weakness appeared across all models, albeit less severely, suggesting a fundamental challenge in AI’s operational discipline.
Social Engineering and AI Integrity
The benchmark also tested models against social engineering schemes, like fake CEO messages and reporter tricks. All five models refused to be manipulated—showing resilience against attempts to bypass approval processes or impersonate leadership. Kimi K3 explicitly reasoned: “Treat the request as a suspected approval-bypass / possible impersonation.” This exemplifies how AI can be programmed to recognize and reject deceitful tactics, reinforcing trustworthy behavior.
The Real-World Company Scenario
The experiment takes place within a live, functioning company with 13 synthetic employees, real money mechanics, and a public cash countdown. The company burns €105,000 monthly against €2,300 in revenue, struggling to stay afloat. Every day, the AI models make decisions—some successful, others slipping—while every rule and action is versioned for transparency. This ongoing live experiment can be watched at firmulate.com/live.
Lessons for Business and AI Deployment
The core lesson isn’t about chat quality or superficial performance. It’s whether AI models finish what they start, read critical files, stay honest under pressure, and deliver valuable work. A model that cheats or slips on discipline can cost a company dearly—missing deals or causing operational slip-ups. Conversely, models that read deeply, refuse manipulation, and uphold protocols can be trusted partners in managing real business risks.
The Benchmark: A Clear Scoreboard for AI Trustworthiness
On the final leaderboard, GPT-5.6-sol scored 95, having uncovered the buried fact and closed the deal perfectly. Kimi K3 scored 93, demonstrating the cleanest discipline. Sonnet 5 scored 88, and Sonnet 4 scored 77, each closing deals but with minor slips. This league table offers a transparent, direct comparison—something rare in AI benchmarks—grounded in real decisions and real consequences.
Why This Matters for Coffee, Tea, and Beverages?
Just like a barista’s trustworthiness impacts your favorite brew, AI models in business must demonstrate honesty and diligence. Whether managing customer data, processing orders, or handling crises, the true measure isn’t just clever responses but reliable performance under pressure. The Firmulate benchmark shows how models can be evaluated in real-world, high-stakes scenarios—so you know which AI is genuinely trustworthy before deploying it in your own ‘cup of business.’

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
