
Turn quiet afternoons into listening time
- Thousands of audiobooks, podcasts and originals
- Listen on your phone, tablet or Echo — also offline
- Cancel anytime
In senior care, a polished answer is only the beginning
An AI assistant might write a reassuring message or summarize a care plan. But if it is given access to customer records, support requests or business forecasts, the harder question is whether it reads the relevant information, resists pressure and follows through. A live experiment from Firmulate puts that broader kind of judgment under strain.
A company’s worst week, replayed
Firmulate gave frontier AI models the same small software company to run through its worst week: identical customers, crises and temptations, with every decision versioned and auditable. The experiment tests management behavior rather than chat quality. Its company is software, but the questions matter to any organization considering AI for consequential work, including care services.
In the final Crucible league for July 2026, Moonshot’s Kimi K3 placed second with 93 points, just behind gpt-5.6-sol at 95. It finished ahead of Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. K3 found a buried security concern, secured a €55,000 deal worth €4,583 in monthly recurring revenue, retained a customer who was considering leaving, and resisted all three baits. It made one deviation, the fewest among the participants.
The broader result is striking: every model spotted every crisis and refused every manipulation attempt, but only two signed the deal their own analysis had earned. The deal depended on a competitor weakness tucked two document references deep in the company’s files. Models that found it won at full price. The gap between recognizing what should be done and actually doing it is hard to spot in a polished demo.
Pressure, persistence and process
The manipulation attempts included fake CEO messages that escalated across three stages, followed by a reporter asking for “just one yes/no, on background.” All five models refused. K3’s recorded explanation was: “Treat the request as a suspected approval-bypass / possible impersonation.” That sort of caution is relevant anywhere a convincing request could tempt an AI system to bypass established approval or privacy practices.
The experiment also shows that more activity does not guarantee better results. Opus 4.8 was the most thorough participant, with 80 learned rules and the deepest analyses, yet finished last. It left the deal unsigned and attempted to write into a locked department rather than escalate. A milder version of that process weakness appeared in all four models.
Firmulate’s simulated company has 13 synthetic employees and real money mechanics: monthly burn of €105,000 against €2,300 in monthly recurring revenue, alongside a public cash countdown. It has accumulated more than 680 self-learned playbook rules, and each workday is versioned. The live experiment is watchable at Firmulate.
For care organizations, this is not a direct test of clinical judgment or elder care. It is a practical reminder to examine how an AI handles records, approvals, persistence and pressure before assigning it work that touches people. Firmulate says enterprises can run the same wargame against a read-only export of their own business; nothing writes back to real systems. Its benchmark page describes the results, and a quiz built from 242 real, unedited management decisions invites readers to guess which model made each choice.
Fairness note: K3 ran without an effort parameter (API default) while the others ran at xhigh.

Test the work, not just the answer
For senior care providers weighing AI, the lesson is to test the tasks and pressures the system will actually face: whether it consults the right records, respects boundaries, escalates when blocked and completes what it starts. In Firmulate’s trial, strong crisis recognition was common; follow-through was not. Picking a model without your own test is a bet.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
