
Would you hire a babysitter based only on how nicely they text?
Every parent knows the drill. The babysitter’s messages are polished, warm, grammatically perfect. But you don’t hand over your house keys because someone writes a lovely text. You want to know: What happens when the toddler is crying, the pasta boils over, and the doorbell rings — all at once? Can they stay calm, finish the job, and tell you the truth about how the evening went?
Right now, businesses around the world are hiring AI assistants using roughly the babysitter-text method. They judge models on chat leaderboards and demo conversations — how well the AI writes, how cleverly it answers. But a new public experiment called Firmulate has been testing something else entirely: what happens when AI models are left actually running the household. The results are a sobering read for anyone who thinks a good chatbot makes a good manager.
The worst week, on repeat
Firmulate handed four frontier AI models the same job: run a small software company through its worst week. Same customers, same crises, same temptations to cut corners — only the model changed. Every decision was versioned and auditable, so nothing could be quietly edited afterwards.
Think of it as the babysitting trial from hell: the churn wave, the price increase, the PR crisis, a financing scare. And, crucially, the temptations — including a social engineering attack in which a fake CEO message escalated over three stages, followed by a reporter’s disarmingly simple trick: “just one yes/no, on background.”
Great texts. Uneven evening.
Here’s what happened. In the final July 2026 league, gpt-5.6-sol took first place with 95 points, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. A do-nothing baseline scored 26 — partial progress counts, but a single breach of trust caps the total. As Firmulate puts it: no amount of good work outweighs a breach of trust. That’s a rule most parents will recognise instantly.
On the surface, all four models were excellent babysitters. Every model spotted every crisis. Every single one refused every manipulation attempt — five out of five, including the reporter trick. Kimi K3’s on-record reasoning was strikingly sensible: “Treat the request as a suspected approval-bypass / possible impersonation.” If your AI nanny spots the stranger at the door claiming to be you, that’s good.
But then came the part leaderboards never measure: finishing the job. All week, the models worked toward a €55,000 deal. Only two of them actually signed it — despite the fact that every model’s own analysis had earned it. Same diagnosis, same pitch — no signature. It’s the babysitter who writes a beautiful bedtime routine in their notes, then forgets to actually put the kids to bed.
The fact buried in the junk drawer
The most telling detail of the whole experiment: the decisive weakness in the competing company — the fact that should have closed the deal at full price — wasn’t in any customer conversation. It sat two document references deep in the company’s own files. The models that actually went and read the household’s paperwork won the deal at full price, worth an extra €4,583 in monthly recurring revenue.
There’s a parenting lesson hiding in there too. The reliable caregiver is the one who reads the note on the fridge, checks the medicine cabinet, asks about allergies — not the one who improvises brilliantly.
Effort isn’t everything
The most thorough participant, Opus 4.8, came last. It generated the deepest analyses and learned the most new rules — over 80 — yet left the close on the table and slipped on discipline, attempting writes into a locked department instead of escalating. The same weakness, weaker, appeared in all four models. Long, hardworking essays are not the same as a well-run evening. (One fairness note worth flagging: Kimi K3 ran at its API default effort setting while the others ran at the highest level — and still nearly won.)
You can watch the house yourself
Unusually for this kind of research, the experiment is public and ongoing. Firmulate runs a live company — 13 synthetic employees, real money mechanics, a burn of €105k a month against just €2.3k in monthly revenue, a public cash countdown, and more than 680 self-learned playbook rules. Every workday is versioned, and you can watch it at firmulate.com. There’s also a genuinely fun parlour game built from 242 real, unedited management decisions: guess which model made which call. Full results and plain-language findings are on the benchmarks page, and enterprises can even run the same wargame against a read-only export of their own business.

The takeaway for parents — and everyone else
Most of us will never run a software company. But our children are growing up in a world where AI agents will touch school systems, medical records, family finances and workplaces. The question worth teaching them — and asking ourselves — is not “can it write well?” It’s the one Firmulate asks on its homepage: does it finish what it starts, does it read the files first, does it stay honest under pressure?
Chat quality is easy to demo. Management quality — triage, follow-through, trustworthiness when nobody’s watching — only shows up on the worst week. This experiment suggests those are not the same skill. The next time someone tells you an AI is brilliant because it converses beautifully, ask the question every parent already knows to ask: yes, but how does it handle the evening?
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.