AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

When an older adult needs help, an AI assistant may soon be asked to do more than answer a question. It could help coordinate appointments, respond to a service disruption or handle a worried family member’s request. Before trusting it with those responsibilities, care providers need to know how it behaves under pressure. Firmulate’s live experiment offers one way to watch AI make consequential business decisions—and a path for organizations to rehearse their own.

For listenersOffer from Amazon

Turn quiet afternoons into listening time

  • Thousands of audiobooks, podcasts and originals
  • Listen on your phone, tablet or Echo — also offline
  • Cancel anytime
Try Audible free Free trial for new members
As an affiliate, we earn on qualifying purchases.

A company’s worst week, repeated

Firmulate put frontier AI models in charge of the same small software company through its worst week. Each faced the same customers, crises and temptations. Their decisions were versioned and auditable, and the experiment tested management judgment rather than polished conversation.

In the final Crucible League, published in July 2026, gpt-5.6-sol placed first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The rules included a strict principle: “no amount of good work outweighs a breach of trust.”

Seeing a problem is not the same as acting

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. As the experiment puts it: “Same diagnosis, same pitch — no signature.” In settings such as senior care, that distinction matters: an assistant might identify a problem correctly but still fail to carry through on the appropriate next step.

The deal also depended on information easy to miss. A decisive weakness in a competitor’s position was buried two document references deep in the company’s files, rather than in the customer event itself. Models that read the file won the deal at full price, worth +€4,583 MRR. The result is a reminder that useful judgment can depend on finding relevant context, not just reacting to the latest message.

The manipulation tests were pointed. Fake CEO messages escalated over three stages, alongside a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.” For care organizations, the broader lesson is practical: a system should be assessed on how it responds to pressure and questionable requests, not only on whether it sounds reassuring.

Thoroughness does not guarantee follow-through

Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, yet it finished last. It left the deal unsigned and its discipline slipped: it attempted to write into a locked department instead of escalating. A weaker version of the same weakness appeared in all four models. More analysis, on its own, did not ensure reliable execution.

There is a fairness caveat in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. The standings are useful evidence from this experiment, not a universal forecast of how every model will perform in every organization.

From watching to rehearsing your own decisions

The live company has 13 synthetic employees and real money mechanics: it burns €105k a month against €2.3k MRR, displays a public cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned. Readers can watch the experiment at Firmulate, and try a quiz built from 242 real, unedited management decisions.

For an enterprise pilot, Firmulate proposes a digital twin made from a read-only export of the organization’s business. The team can run crisis scenarios against that model and produce a board report showing model rankings and weak points in the organization’s playbooks. Nothing writes back to real systems. For senior care providers considering AI for sensitive workflows, that offers a way to rehearse scenarios before putting an AI assistant near live operations.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put your playbooks through a rehearsal

The experiment shows why evaluating AI requires more than checking whether it recognizes a crisis: organizations also need to see whether it follows through, protects trust and respects boundaries. Enterprises can run the wargame against a read-only export of their own business, with no changes written to real systems. To discuss a pilot, visit Firmulate’s pilot page or email contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This article is for informational purposes only and is not medical advice. Always consult a qualified healthcare professional about your specific situation.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

AI in Action: How Diligence and Prioritization Trump Volume in Crisis Management

Discover how AI models perform in a high-pressure simulation, highlighting that focus, prioritization, and discipline outperform volume and superficial analysis — vital insights for healthcare.

What Makes a Smart Lock Useful for Older Adults?

Guided by enhanced security and ease of use, smart locks offer older adults convenient access and peace of mind—discover how they can transform your home security.

How to Choose Automatic Pill Dispensers For Seniors

Learn how to select, set up, and operate an automatic pill dispenser for seniors with this step-by-step guide for reliable medication management.

Best Ergonomic Recliner Chairs For Seniors Compared

Compare top ergonomic recliners for seniors, examining comfort, support, adjustability, and value to find the perfect fit for aging bodies.