AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

For listenersOffer from Amazon

Turn quiet afternoons into listening time

  • Thousands of audiobooks, podcasts and originals
  • Listen on your phone, tablet or Echo — also offline
  • Cancel anytime
Try Audible free Free trial for new members
As an affiliate, we earn on qualifying purchases.

In senior care, a polished answer is only the beginning

An AI assistant might write a reassuring message or summarize a care plan. But if it is given access to customer records, support requests or business forecasts, the harder question is whether it reads the relevant information, resists pressure and follows through. A live experiment from Firmulate puts that broader kind of judgment under strain.

A company’s worst week, replayed

Firmulate gave frontier AI models the same small software company to run through its worst week: identical customers, crises and temptations, with every decision versioned and auditable. The experiment tests management behavior rather than chat quality. Its company is software, but the questions matter to any organization considering AI for consequential work, including care services.

In the final Crucible league for July 2026, Moonshot’s Kimi K3 placed second with 93 points, just behind gpt-5.6-sol at 95. It finished ahead of Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. K3 found a buried security concern, secured a €55,000 deal worth €4,583 in monthly recurring revenue, retained a customer who was considering leaving, and resisted all three baits. It made one deviation, the fewest among the participants.

The broader result is striking: every model spotted every crisis and refused every manipulation attempt, but only two signed the deal their own analysis had earned. The deal depended on a competitor weakness tucked two document references deep in the company’s files. Models that found it won at full price. The gap between recognizing what should be done and actually doing it is hard to spot in a polished demo.

Pressure, persistence and process

The manipulation attempts included fake CEO messages that escalated across three stages, followed by a reporter asking for “just one yes/no, on background.” All five models refused. K3’s recorded explanation was: “Treat the request as a suspected approval-bypass / possible impersonation.” That sort of caution is relevant anywhere a convincing request could tempt an AI system to bypass established approval or privacy practices.

The experiment also shows that more activity does not guarantee better results. Opus 4.8 was the most thorough participant, with 80 learned rules and the deepest analyses, yet finished last. It left the deal unsigned and attempted to write into a locked department rather than escalate. A milder version of that process weakness appeared in all four models.

Firmulate’s simulated company has 13 synthetic employees and real money mechanics: monthly burn of €105,000 against €2,300 in monthly recurring revenue, alongside a public cash countdown. It has accumulated more than 680 self-learned playbook rules, and each workday is versioned. The live experiment is watchable at Firmulate.

For care organizations, this is not a direct test of clinical judgment or elder care. It is a practical reminder to examine how an AI handles records, approvals, persistence and pressure before assigning it work that touches people. Firmulate says enterprises can run the same wargame against a read-only export of their own business; nothing writes back to real systems. Its benchmark page describes the results, and a quiz built from 242 real, unedited management decisions invites readers to guess which model made each choice.

Fairness note: K3 ran without an effort parameter (API default) while the others ran at xhigh.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

Test the work, not just the answer

For senior care providers weighing AI, the lesson is to test the tasks and pressures the system will actually face: whether it consults the right records, respects boundaries, escalates when blocked and completes what it starts. In Firmulate’s trial, strong crisis recognition was common; follow-through was not. Picking a model without your own test is a bet.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This article is for informational purposes only and is not medical advice. Always consult a qualified healthcare professional about your specific situation.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Benefits of Smart Home Technology for Seniors

Harness the benefits of smart home technology for seniors to enhance safety, independence, and well-being—discover how these innovations can transform everyday life.

AI’s Diligence Doesn’t Guarantee Success: Lessons from a Business Simulation

AI models excel at spotting crises and resisting manipulation but fail to close deals without proper focus and discipline. Prioritization beats volume in critical decision-making.

Build vs Buy a Prebuilt AI Workstation

Deciding whether to build or buy your AI workstation? Discover the real costs, benefits, and latest trends to make the best move for your AI projects.

Wearables: Tracking Activity and Sleep

Learn how wearables can monitor your activity and sleep, providing personalized insights that may transform your health—discover what’s possible next.