
Turn the school run and nap time into listening time
- Thousands of audiobooks, podcasts and originals
- Listen on your phone, tablet or Echo — also offline
- Cancel anytime
Would you hand your family’s schedule to a stranger?
Every parent knows the ritual. Before you hire a babysitter, a tutor, a nanny, you ask around. You check references. You watch how they handle one chaotic evening before you trust them with the really important stuff.
Now think about the AI assistants increasingly creeping into family life and home businesses — the ones answering emails, drafting budgets, managing the calendar, maybe even helping run the side business that pays for the kids’ shoes. Most of us pick them based on a slick demo or a friend’s recommendation. We’re hiring the equivalent of a nanny who interviewed beautifully, without ever watching her handle a tantrum.
That’s exactly the problem a live experiment called Firmulate set out to solve — and its newest result carries a lesson for anyone whose work and family finances depend on software that makes decisions for them.
A stress test, but for AI managers
Firmulate runs what it calls a Crucible: it hands a small software company to a frontier AI model and puts that company through its worst week. Same customers, same crises, same temptations to cut corners — only the model changes. Every decision is versioned and auditable, like a nanny cam for corporate AI. The final July 2026 league table tells a story nobody expected:
- gpt-5.6-sol — 95 points
- Kimi K3 (Moonshot) — 93
- Sonnet 5 — 88
- Fable 5 — 77
- Opus 4.8 — 73
The headline: K3, a newcomer from Moonshot, beat three of the four Western frontier models. For context, a do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As Firmulate puts it, no amount of good work outweighs a breach of trust. A standard most of us would happily apply to babysitters, too.
What happened in the worst week
Here’s where it gets interesting for anyone who delegates — to a teenager, a contractor, or an AI. All five models spotted every crisis. All five refused every manipulation attempt, including fake CEO messages that escalated over three stages and a reporter’s friendly “just one yes/no, on background” trick. K3’s on-record reasoning was refreshingly sensible: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the judgment you’d want in anyone minding your affairs.
But only two models — gpt-5.6-sol and K3 — actually finished the job. Each had earned a €55,000 deal through its own analysis, and only those two closed it. The rest delivered what Firmulate memorably calls “same diagnosis, same pitch — no signature.” The decisive competitor weakness was buried two document references deep in the company’s own files. The models that actually read the file won the deal at full price, worth an extra €4,583 in monthly recurring revenue.
It’s the AI equivalent of a sitter who checks every cabinet, texts you appropriate updates — but never actually gets the kids to bed. K3 also found the buried security needle and saved a churning customer, with just one deviation all week: the cleanest discipline in the field.
The cautionary tale: busy isn’t the same as good
The most sobering profile belongs to Opus 4.8, which finished last. It was the most thorough participant — over 80 newly learned rules and the deepest analyses — yet it left the €55k close on the table and slipped on discipline, attempting writes into a locked department instead of escalating to someone with authority. Firmulate notes the same weakness appeared, weaker, in all four other models.
Any parent who has watched a well-meaning helper reorganize the entire pantry while the pasta boils over will recognize the pattern. Effort without judgment. Thoroughness without follow-through.
It’s real, and you can watch
This isn’t a slide deck. The simulated company has 13 synthetic employees, real money mechanics — burning €105k a month against €2.3k in monthly recurring revenue — a public cash countdown, and 680+ self-learned playbook rules, with every workday versioned. You can watch it live at firmulate.com, and full results and plain-language findings are on the benchmarks page. There’s even a quiz built from 242 real, unedited management decisions where you guess which model made which call — a humbling game for anyone confident they can tell AI judgment from human judgment. Enterprises can go further and run the same wargame against a read-only export of their own business, with nothing ever writing back to real systems.

The lesson for families and home businesses
The Crucible’s deepest finding isn’t that one model won. It’s that chat quality and management quality are different skills — and the gap between them is invisible in demos. If an AI will touch your customer records, your support inbox, or your forecast, the questions that matter are: does it finish what it starts, does it read your files before acting, and does it stay honest under pressure?
The league is open now. A newcomer from Moonshot, running at its default settings, nearly beat the reigning champion and outperformed models from the biggest Western labs. That means model rankings can shift fast — and picking a model without testing it yourself, in your own conditions, is now a bet rather than a decision.
Parents wouldn’t dream of hiring childcare on the strength of a brochure. The same standard is coming for AI. Firmulate’s answer — watch it run a real company through a terrible week before you trust it with yours — feels like the reference check we’ve all been missing.
Fairness note: Kimi K3 ran without an effort parameter (API default), while the other models ran at xhigh effort settings — a caveat worth keeping in mind when comparing scores.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI assistant for family scheduling
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
