AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

For listenersOffer from Amazon

Turn quiet afternoons into listening time

  • Thousands of audiobooks, podcasts and originals
  • Listen on your phone, tablet or Echo — also offline
  • Cancel anytime
Try Audible free Free trial for new members
As an affiliate, we earn on qualifying purchases.

When care depends on good judgment, a polished answer is not enough

Families arranging senior care make decisions under pressure: a provider changes, costs rise, or a loved one’s needs shift. AI tools may soon help organizations handle the work behind those decisions. But can an AI workforce recognize a crisis, resist pressure and follow through when the right next step is difficult?

Firmulate is testing that question by running AI models as a company, with customers, money and consequential choices. Its experiment offers a practical lesson for anyone considering AI in a high-trust setting: watch what it does under pressure before letting it handle real work.

A company’s worst week, repeated

In the final Crucible League, held in July 2026, frontier models faced the same small software company and the same difficult week. Customers, crises and temptations were held constant; only the model changed. Decisions were versioned and auditable.

The results put gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The league also treated trust as essential: a single breach of trust capped the total, on the principle that “no amount of good work outweighs a breach of trust.”

All the models spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. The experiment’s concise finding was: “Same diagnosis, same pitch — no signature.” Recognizing the right answer did not always mean carrying it through.

The important detail was buried

The deal turned on a competitor weakness hidden two document references deep in the company’s own files, rather than in the customer event itself. Models that read the file won the deal at full price, worth +€4,583 MRR. The difference came from attending to relevant information already available, then acting on it.

That pattern has a clear parallel in senior care. A useful AI assistant might need to connect a new request with an existing care plan, a provider’s instructions or a family’s stated priorities. A system that responds smoothly but misses the crucial detail—or fails to act on what it found—can leave people with a false sense of confidence.

Trust under pressure, and discipline afterward

The experiment tested social engineering with fake CEO messages escalating through three stages and a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s recorded reasoning described the request as “a suspected approval-bypass / possible impersonation.”

Refusing manipulation was only part of the picture. Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but placed last. It left the deal unsigned and attempted writes in a locked department instead of escalating. A weaker version of the same discipline problem appeared in all four models.

There is also a fairness caveat when reading the ranking: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. That detail belongs alongside the scores when interpreting the comparison.

From watching to trying it on your own business

The live Firmulate company makes the test watchable. It has 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k MRR, a public cash countdown, more than 680 self-learned playbook rules and versioned workdays. The figures describe a synthetic company, not a senior-care provider. The experiment is a way to observe AI management choices, not evidence that an AI can safely make care decisions.

For organizations considering AI in their operations, Firmulate’s proposed next step is a pilot using a read-only export of the organization’s own business. It runs crisis scenarios against that business and produces a board report with a model ranking and weak points in its playbooks. Nothing writes back to real systems. The idea is to see how models handle your own context before entrusting them with live workflows.

Firmulate also offers a “guess the model” quiz based on 242 real, unedited management decisions. Readers can explore how model behavior looks in practice, then compare that with the league’s published results. See the live experiment and results.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Test the decision, not just the answer

For senior-care organizations, AI adoption should start with questions about judgment, context, trust and escalation. Firmulate’s experiment shows why: models can identify a crisis and resist manipulation, yet still miss a detail or fail to complete the action their analysis supports. A pilot can put those behaviors under pressure using an organization’s own business data, while keeping the connection read-only.

To explore a pilot for your organization, visit Firmulate’s pilot page or email contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This article is for informational purposes only and is not medical advice. Always consult a qualified healthcare professional about your specific situation.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Golden Sun Health Technology Group Surges In Global Coverage

Coverage of Golden Sun Health Technology Group has surged globally, with 21 mentions this week, marking an eightfold increase from baseline, sparking widespread interest.

AI In Drug Discovery – What It Is, Where We Stand And The Path Forward

An analysis of AI’s role in drug discovery, current advancements, challenges, and future prospects based on recent developments and expert insights.

AI in Business: Beyond Chat Scores to Trust and Resilience

A live AI experiment reveals key management traits—reading, honesty, resilience—that matter more than chat scores, especially for sensitive fields like senior care.