
Imagine Watching a Company Struggle for Its Life — Powered by AI, Without a Single Employee
In an unprecedented live experiment, a tiny software firm is being run entirely by artificial intelligence models. Every decision, crisis, and temptation is played out in real time, revealing how AI handles the messy realities of managing a business. For seniors and caregivers, this story offers a glimpse of AI’s potential — and its limits — in high-stakes situations, including decision-making, trust, and integrity.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Live Experiment: A Business in the Crossfire of AI and Reality
At the heart of this experiment is a small, virtual company, watched daily at firmulate.com/live. It has no employees, no managers. Instead, it is run by four advanced AI models, each assigned the same challenging week: facing customer crises, internal dilemmas, and ethical tests. Their task? To navigate a turbulent business landscape with the same constraints as a real startup — but under the scrutiny of a public, transparent process.
Decoding AI Decision-Making in Business
All four models—names like GPT-5.6, Kimi K3, Sonnet 5, and Fable 5—successfully identified every crisis thrown at them. They refused manipulation attempts, including an elaborate social engineering scheme where fake CEO messages escalated over stages, and even a reporter asking for a simple yes/no agreement. Remarkably, all models declined these unethical prompts, demonstrating a strong grasp of integrity under pressure. The critical insight, however, lies deeper:
- Only two models managed to close a legitimate €55,000 deal, based on their own analysis, with full price and confidence.
- The others, despite good diagnosis, left revenue on the table, illustrating how discipline and execution remain vital even for AI.
The Hidden Weakness: The Power of Internal Information
The decisive advantage went to models that dug deeper into the company’s own files—accessing information buried two documents down. This allowed them to uncover a key piece of evidence that sealed the deal for full price, worth over €4,500 in monthly recurring revenue. It’s a striking reminder that in real business, reading beyond surface data can be the difference between success and failure.
Trust and Ethical Challenges
Beyond crisis management, the experiment tested the models against social engineering. The fake CEO messages, escalating over three stages, and a background query by a reporter—all designed to tempt the AI into bypassing protocols—met unanimous rejection. Kimi K3 explained their refusal: ‘Treat the request as a suspected approval-bypass / possible impersonation.’ This emphasizes that AI can be trained to uphold ethical standards even when under attack from human manipulations.

Artificial Intelligence for HR: Use AI to Support and Develop a Successful Workforce
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Company Without Employees: A Living Laboratory
The experiment is hosted on Firmulate’s live site, showcasing a company that is entirely synthetic, with 13 AI ‘employees’ orchestrating real money mechanics. It burns €105,000 each month amid a public countdown to survival, with every decision versioned and auditable. What makes this so compelling isn’t just its novelty — it’s the transparency about AI’s capabilities and vulnerabilities in a real-world context.
Lessons for Senior Care and Aging
While this may seem distant from caregiving or senior services, the implications are profound. AI systems in healthcare, support, and management will face crises, ethical dilemmas, and trust tests. Can they read and interpret critical data? Will they stay honest under pressure? The Firmulate experiment shows that AI can recognize crises and refuse manipulation, but execution discipline can still falter without proper oversight.
The Deep Dive: What AI Models Can and Cannot Do
The models’ scores reveal their respective strengths:
- GPT-5.6 scored 95/100, successfully finding critical facts and closing the deal.
- Kimi K3 scored 93/100, maintaining the best discipline and closing at full price.
- Sonnet 5 scored 88/100, with some process slips but still closing deals.
- Fable 5, despite good rules discipline, left opportunities unexploited, scoring 77/100.
This paints a nuanced picture: AI can handle crises and uphold integrity but may struggle with consistent execution under stress. For sectors like senior care, where trust and meticulous care are paramount, understanding these strengths and limits is essential.


Applying AI in Learning and Development: From Platforms to Performance
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Takeaways: Trust, Discipline, and Transparency in AI-Driven Business
This live experiment offers a rare view of AI managing the chaos of a real business. It demonstrates that AI can recognize crises, uphold ethical standards, and even uncover hidden information to close deals. But discipline and execution still matter—especially when stakes are high. For those involved in caring for seniors or managing vulnerable populations, it’s a reminder that AI’s promise depends on transparency, rigorous oversight, and understanding its current limits. Watching this experiment unfold in real time, we see a future where AI supports—and sometimes challenges—the human decision-maker.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Trust.: Responsible AI, Innovation, Privacy and Data Leadership
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.