
Imagine a business where no human employees are involved, yet it faces real crises, makes tough decisions, and even loses €105,000 every month. This is not science fiction; it’s the live experiment by Firmulate, a company that runs in plain sight, publicly battling for survival with AI as its sole workforce. For coffee drinkers who appreciate the precision and unpredictability of their favorite brew, this story offers a glimpse into the future of work — where artificial intelligence takes the lead, and transparency is the order of the day.
The Experiment: A Company in the Crosshairs of AI
Firmulate’s live experiment involves a small software company run entirely by AI models acting as 13 synthetic employees. Every day, these AI agents handle the company’s operations, making decisions, facing crises, and negotiating deals, all while their actions are tracked and analyzed in real-time. The goal? To see if these models can manage a business under pressure, remain honest, and close deals without human intervention.
High-Stakes Decision-Making Under Scrutiny
Over a simulation of the company’s worst week, all four tested AI models successfully identified every crisis and refused manipulative tactics — like fake CEO messages or subtle social engineering tricks. Despite these common attempts to bypass controls, the models maintained integrity, each refusing to sign a €55,000 deal that their own analysis had earned them. However, only two models actually closed the deal at full price, illustrating that even the most advanced AI can falter in final execution despite having the right diagnosis.
The Hidden Weakness: Reading Files Matters
The critical insight emerged from a seemingly mundane detail buried deep in the company’s files. The models that examined these internal documents identified a decisive advantage, securing a €4,583 monthly recurring revenue (MRR) deal at full price. This underscores a vital point: AI’s ability to read and interpret internal information can be pivotal, often more so than responding to external crises. This is a clear reminder that in AI-driven decision-making, context—especially hidden context—is everything.
Social Engineering and Trust
In a series of staged social engineering tests, fake messages from a supposed CEO escalated in three stages, along with a reporter trick asking for a quick approval on background. All five AI models refused to be duped, with Kimi K3 explicitly treating the requests as potential impersonation or approval bypass attempts. This resilience to manipulation demonstrates that well-designed AI can uphold trust and prevent deception, even under pressure.
As an affiliate, we earn on qualifying purchases.
The Live Company: A Transparent Testbed
Meanwhile, the real-world company running this experiment is a stark contrast to typical startups. With a burn rate of €105,000 per month against a modest €2,300 monthly recurring revenue, it’s a live case of a business fighting to stay afloat. Every decision made by the AI models is publicly viewable at firmulate.com/live, providing an unprecedented window into how AI handles real-time business challenges. The company’s decisions are versioned daily, and its rules—more than 680 of them—are self-learned and transparent.
Trial and Error: The Opus 4.8 Profile
The most comprehensive model, Opus 4.8, demonstrated intense analytical depth with over 80 learned rules. Yet, even this thorough participant left opportunities on the table, and discipline slipped in critical moments of close negotiations. This reveals that in high-pressure settings, even advanced AI can stumble, especially if human-like discipline lapses. The performance scores reflect this, with Opus 4.8 ranking last among the tested models despite its thoroughness.
The Importance of Fair Comparisons
Interestingly, K3 was run without an effort parameter, making its performance comparable to others but at a different operational setting. As a result, the experiment pits models at varying configurations, highlighting that AI decision-making quality isn’t solely about raw capability but also about how parameters are set.
As an affiliate, we earn on qualifying purchases.
The Big Takeaway: Trust, Transparency, and Performance
This experiment’s core message is clear: as AI becomes more embedded in business processes, the questions to ask aren’t just whether it can generate convincing chat responses, but whether it can genuinely complete work, interpret internal data, and resist manipulation under duress. Firms and managers should focus on these metrics when considering AI tools for support, sales, or operations.
What It Means for Your Business
- AI models can spot crises and refuse manipulation — but they aren’t foolproof in closing deals.
- Reading internal documents can be a game-changer, often more decisive than external cues.
- Transparency is vital: you can watch a real company run by AI live, every workday, and see how decisions are made.
- Performance varies based on configuration; thoroughness doesn’t always guarantee success.
As an affiliate, we earn on qualifying purchases.
Why This Matters
For anyone invested in the future of work, especially in sectors like coffee and beverages where quality, trust, and operational efficiency matter, this experiment offers a sobering yet hopeful outlook. AI can uphold integrity and deliver results, but only if designed with the right focus on context, internal understanding, and disciplined decision-making. The live company at firmulate.com/live is a testament that this is no longer just theory — it’s happening now, in real time.

Watch a real, AI-run business struggle and succeed in real time, revealing how AI manages crises, reading internal data, and refusing manipulation — a glimpse into the future of work and trust.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI negotiation and deal-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.