
Imagine a company with no human employees, losing €105,000 every month, yet still working every business day, battling crises and making decisions in real time. For senior care and aging sectors, where trust and reliability are everything, this experiment offers a rare glimpse into the future of AI management — one where transparency and honesty are put to the test every single day.
The Live Experiment: An AI Company Under Pressure
Firmulate has created a groundbreaking live demonstration — a company run entirely by artificial intelligence models, but with real financial mechanics, crises, and decision-making. It features 13 synthetic employees, each guided by thousands of self-learned rules, operating in a real-time environment accessible at firmulate.com/live. Each workday, the company faces simulated customer issues, internal crises, and ethical temptations, all designed to test whether AI can truly manage complex, high-stakes business operations.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Critical Findings: Honesty, Detection, and Decision-Making
The experiment tested four advanced AI models, including GPT-5.6, Kimi K3, Sonnet 5, and Opus 4.8. All models successfully identified every crisis presented to them and refused manipulation attempts, such as fake CEO messages or covert offers. This suggests that AI can be remarkably honest and resistant to social engineering — a vital trait for customer trust and regulatory compliance in senior care or finance sectors.
However, the models’ ability to close deals varied significantly. While GPT-5.6 and Kimi K3 both signed a €55,000 contract, the others did not. The key difference? The models that read deeper into internal company files, unearthing hidden information, secured the full deal. For example, Kimi K3, the newcomer, closed the deal by referencing buried data that others overlooked, earning an additional €4,583 in monthly recurring revenue (MRR). This highlights a crucial point: effective AI decision-making depends not just on surface-level analysis but on deep, context-rich comprehension.

Artificial Intelligence for Cybersecurity: Develop AI approaches to solve cybersecurity problems in your organization
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Challenge of Financial Sustainability
Despite their decision-making capabilities, the company struggles financially, burning through €105,000 each month against only €2,300 in MRR. The public cash countdown underscores the fragility of this operation, illustrating the high stakes for AI management systems in real-world applications. Senior care organizations, which rely heavily on integrity and accuracy, should note that AI can make honest, crisis-aware decisions, but financial sustainability and thorough internal comprehension remain critical hurdles.

How AI Agents Work: Tools, Memory, and Autonomous Decision-Making (The AI Security & Hacking Bible: Protect and Exploit LLMs and Autonomous Agents)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Built-in-Public and Transparent Progress
This experiment is built in public, with every decision, version, and rule openly documented and auditable. The company updates itself twice daily, and every workday’s performance is available for scrutiny. The process offers a rare, transparent view of how AI models behave in a complex, real-time environment — information that could be invaluable for organizations considering AI integration in sensitive fields like healthcare or elder support.

Building AI-Powered Financial Products: Use responsible AI to launch ROI-driven FinTech products at scale
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Lessons for the Future of AI Management
The Firmulate experiment demonstrates that AI can be honest under pressure, detect social engineering attempts, and even find hidden information to close deals. Yet, it also reveals ongoing challenges: financial viability and disciplined escalation of issues need improvement. The most thorough model, Opus 4.8, missed some opportunities due to discipline lapses, showing that even advanced AI models require careful tuning and oversight.
For sectors that depend on trust, accuracy, and ethical decision-making, these insights are crucial. While AI is not yet fully autonomous or financially sustainable at scale, its capacity for honesty and crisis detection marks a significant step forward.

The live AI company at firmulate.com/live offers a transparent, real-time look into how artificial intelligence can manage complex business operations. It proves AI can be honest, identify hidden risks, and make crucial decisions — vital qualities for senior care, healthcare, and finance sectors aiming for trustworthy automation. But it also highlights the ongoing challenges of financial sustainability and disciplined escalation, reminding us that AI’s future in management is as much about oversight as it is about innovation.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html