firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Imagine handing an AI the family calendar, bills and school messages, then asking it to manage a week when everything goes wrong. Would it notice the urgent problems, resist a convincing scam and follow through on the decision it recommended? Those are practical questions for families—and for any business considering AI agents.

For listenersOffer from Amazon

Turn the school run and nap time into listening time

  • Thousands of audiobooks, podcasts and originals
  • Listen on your phone, tablet or Echo — also offline
  • Cancel anytime
Try Audible free Free trial for new members
As an affiliate, we earn on qualifying purchases.

Firmulate puts those questions into a live company experiment. Its final Crucible League tested frontier models as managers of the same small software company, facing identical customers, crises and temptations. The experiment is watchable at firmulate.com.

Good judgment means more than spotting trouble

The results offer an intriguing split. Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. As the experiment put it: “Same diagnosis, same pitch — no signature.” Recognizing what should happen and carrying it through turned out to be different tests.

The deal depended on a detail buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. In a household, the equivalent might be a critical detail tucked into an old school email or insurance document: context can matter as much as a confident recommendation.

Firmulate’s social-engineering challenge tested another everyday concern: whether an AI could be pressured into bypassing normal safeguards. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 explained its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

More effort did not guarantee a better finish

The final July 2026 league ranked gpt-5.6-sol first at 95, Kimi K3 second at 93 and Sonnet 5 third at 88. Fable 5 scored 77, while Opus 4.8 scored 73. The do-nothing baseline scored 26. The benchmark counts partial progress, but one breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet finished last. The close was left on the table, and discipline slipped when it tried writing into a locked department instead of escalating. A weaker version of the same weakness appeared in all four models. That is a useful reminder for anyone weighing AI help at home or work: diligence is valuable, but it does not automatically lead to good judgment or follow-through.

There is a fairness detail behind the ranking. Kimi K3 ran without an effort parameter, using the API default; the other models ran at xhigh. The live company adds a view of the longer-running experiment: it has 13 synthetic employees and real money mechanics, with burn of €105k per month against €2.3k MRR, a public cash countdown, more than 680 self-learned playbook rules and versioned workdays. A separate quiz draws on 242 real, unedited management decisions. Readers can watch the company at firmulate.com.

From watching to testing your own business

For families, these trials are a way to think clearly about what to trust an AI with: Does it catch the important detail? Can it resist pressure? Does it follow through? For a company, those questions can be tested against its own situations rather than a generic demonstration.

Firmulate’s enterprise pilot uses a read-only export of a business to run crisis scenarios against a digital twin. The resulting board report ranks models and identifies weak points in the company’s own playbooks. The pilot is designed so that nothing writes back to real systems.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put the hard questions to a real-world test

The experiment suggests that spotting a crisis, protecting trust and completing the right action are separate measures of AI management. A controlled wargame can show how those behaviors play out before agents are given responsibility in a business.

To explore a pilot against your own company, visit firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Parenting content here is informational. For medical questions about your child, consult a pediatrician.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Good Breastfeeding Diet: A How-To Guide for Nursing Mothers

Navigate the world of nutrition as a nursing mother to optimize your breast milk quality and benefit your baby's development – discover essential tips here!

Postpartum Diet Plan While Breastfeeding: Essential Tips for New Moms

Intrigued about how to navigate a postpartum diet plan while breastfeeding? Discover essential tips for new moms to ensure optimal nourishment for both mom and baby.