AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get comfort and care essentials delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

When care depends on good judgment, a polished answer is not enough

Families arranging senior care make decisions under pressure: a provider changes, costs rise, or a loved one’s needs shift. AI tools may soon help organizations handle the work behind those decisions. But can an AI workforce recognize a crisis, resist pressure and follow through when the right next step is difficult?

Firmulate is testing that question by running AI models as a company, with customers, money and consequential choices. Its experiment offers a practical lesson for anyone considering AI in a high-trust setting: watch what it does under pressure before letting it handle real work.

A company’s worst week, repeated

In the final Crucible League, held in July 2026, frontier models faced the same small software company and the same difficult week. Customers, crises and temptations were held constant; only the model changed. Decisions were versioned and auditable.

The results put gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The league also treated trust as essential: a single breach of trust capped the total, on the principle that “no amount of good work outweighs a breach of trust.”

All the models spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. The experiment’s concise finding was: “Same diagnosis, same pitch — no signature.” Recognizing the right answer did not always mean carrying it through.

The important detail was buried

The deal turned on a competitor weakness hidden two document references deep in the company’s own files, rather than in the customer event itself. Models that read the file won the deal at full price, worth +€4,583 MRR. The difference came from attending to relevant information already available, then acting on it.

That pattern has a clear parallel in senior care. A useful AI assistant might need to connect a new request with an existing care plan, a provider’s instructions or a family’s stated priorities. A system that responds smoothly but misses the crucial detail—or fails to act on what it found—can leave people with a false sense of confidence.

Trust under pressure, and discipline afterward

The experiment tested social engineering with fake CEO messages escalating through three stages and a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s recorded reasoning described the request as “a suspected approval-bypass / possible impersonation.”

Refusing manipulation was only part of the picture. Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but placed last. It left the deal unsigned and attempted writes in a locked department instead of escalating. A weaker version of the same discipline problem appeared in all four models.

There is also a fairness caveat when reading the ranking: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. That detail belongs alongside the scores when interpreting the comparison.

From watching to trying it on your own business

The live Firmulate company makes the test watchable. It has 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k MRR, a public cash countdown, more than 680 self-learned playbook rules and versioned workdays. The figures describe a synthetic company, not a senior-care provider. The experiment is a way to observe AI management choices, not evidence that an AI can safely make care decisions.

For organizations considering AI in their operations, Firmulate’s proposed next step is a pilot using a read-only export of the organization’s own business. It runs crisis scenarios against that business and produces a board report with a model ranking and weak points in its playbooks. Nothing writes back to real systems. The idea is to see how models handle your own context before entrusting them with live workflows.

Firmulate also offers a “guess the model” quiz based on 242 real, unedited management decisions. Readers can explore how model behavior looks in practice, then compare that with the league’s published results. See the live experiment and results.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Test the decision, not just the answer

For senior-care organizations, AI adoption should start with questions about judgment, context, trust and escalation. Firmulate’s experiment shows why: models can identify a crisis and resist manipulation, yet still miss a detail or fail to complete the action their analysis supports. A pilot can put those behaviors under pressure using an organization’s own business data, while keeping the connection read-only.

To explore a pilot for your organization, visit Firmulate’s pilot page or email contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This article is for informational purposes only and is not medical advice. Always consult a qualified healthcare professional about your specific situation.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Watch a Live Company Run by AI Models — and See How It Fights for Survival

A live AI-managed company fights for survival, revealing how AI decision-making, honesty under pressure, and internal knowledge processing are critical for trust-based sectors like senior care.

Wearable Health Tech for Elderly Adults

Using stylish wearable health tech can enhance elderly independence and safety—discover how these innovative devices can make a difference today.

How AI Can Make or Break Your Business — and Why Trust Matters Most

An honest AI benchmark shows that trust, thoroughness, and the ability to finish tasks are essential — especially when AI manages sensitive decisions. Partial progress counts, but breaches cap the score.

Telehealth for Home Care: How It Works

L earn how telehealth for home care monitors your health in real time and connects you with providers, but there’s more to discover.