
Imagine a business operating entirely without human employees, yet constantly battling to stay afloat—publicly showing its every move, every crisis, and every decision. This is not a sci-fi scenario but the real-time experiment of a company that runs on AI models, with its fate unfolding live for anyone to watch.
The Live Experiment: AI as a Company
In an unprecedented move, a small software firm is running its entire operation through AI-powered models, simulating everything from customer crises to strategic decisions. This experiment is not just a showcase but a rigorous test: four frontier AI models have faced the company’s worst week, confronting the same customers, crises, and temptations.
The setup is straightforward yet extraordinary. Every decision made by the AI models is versioned and auditable, providing a transparent view into their reasoning and choices. The company, with no human employees, is openly fighting for survival, burning €105,000 every month against a modest €2,300 in monthly recurring revenue. The entire process is visible at firmulate.com/live.
AI decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Results: Competence, Honesty, and Weaknesses
Surprisingly, all four AI models identified every crisis and refused every manipulation attempt, demonstrating a high level of integrity and awareness. For example, during a social engineering test—where fake CEO messages escalated over multiple stages—every model declined to participate, citing suspicion of impersonation or approval bypass.
Yet, a subtle and critical weakness emerged that was buried deep in the company’s files—two document references away from immediate visibility. The models that read these files and uncovered this hidden information secured a major deal, adding €4,583 to monthly recurring revenue. The models that missed this opportunity left the deal on the table, demonstrating that reading and understanding internal documents can be decisive.
AI business management tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Performance Scores and the League Table
The models’ performance was scored on a scale from 0 to 100:
- gpt-5.6-sol scored 95 — found the buried fact, closed the deal, and demonstrated complete competence.
- Kimi K3 scored 93 — also closed the deal with the cleanest discipline, but missed the buried insight.
- Sonnet 5 scored 88 — closed the deal but with some process slips.
- Fable 5 scored 77 — maintained the best rule discipline but failed to execute the deal after approval.
This league table reveals that success depends not only on surface-level decision-making but on deep reading and understanding of internal information—something that current chat demos often overlook.
AI document analysis software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Building-in-Public Approach
This experiment is a prime example of “build-in-public.” Every decision, every rule learned—over 680 in total—is publicly accessible, versioned, and auditable. The company’s cash countdown is live, and every weekday iteration of AI decision-making is visible to all. It’s a raw, unfiltered look into how AI models might manage real-world business decisions, with all their flaws and strengths laid bare.
AI transparency and audit tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Implications for Business and AI
For anyone managing or deploying AI in a business context, the key takeaway is clear: the question is not just whether these models can generate convincing text or support tasks but whether they can see the full picture, stay honest under pressure, and complete what they start. The company’s experience underscores that reading internal documents, resisting manipulation, and disciplined execution are the real tests.
Moreover, the live site offers enterprises a chance to run their own “wargames”—simulating crises, decision points, and manipulations—without risking actual systems. This is managed entirely through read-only exports, making it a safe environment to evaluate AI readiness before real-world deployment.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html