
Every parent knows the kid: the one who colors inside every line, does the extra-credit worksheet nobody asked for, reads the whole textbook — and then forgets to hand in the actual assignment. Report card says brilliant. Grade says otherwise. It’s one of the hardest lessons of raising children: diligence is not the same as impact. Effort without follow-through doesn’t close the loop.
Turn the school run and nap time into listening time
- Thousands of audiobooks, podcasts and originals
- Listen on your phone, tablet or Echo — also offline
- Cancel anytime
It turns out artificial intelligence learns that lesson the hard way, too — in public, with real money mechanics, and a scoreboard.
The Wargame Where AI Runs a Company
On the public experiment at Firmulate, four frontier AI models were each handed the same job: run a small software company through its worst week. Same customers, same crises, same temptations to cut corners — only the model changed. Every decision was versioned and auditable, so nothing could be hand-waved after the fact.
The final Crucible League standings from July 2026 tell a story any parent of a perfectionist will recognize:
- 1. gpt-5.6-sol — 95 points
- 2. Kimi K3 — 93
- 3. Sonnet 5 — 88
- 4. Fable 5 — 77
- 5. Opus 4.8 — 73
For context, a do-nothing baseline scores 26. Partial progress counts — but a single breach of trust caps the entire total. As the experiment’s own rule puts it: “no amount of good work outweighs a breach of trust.” Which, if you think about it, is also a pretty solid family value.
As an affiliate, we earn on qualifying purchases.
Everyone Saw the Storm; Only Some Closed the Umbrella Deal
Here’s where it gets interesting. All four models spotted every crisis and refused every manipulation attempt — including fake CEO messages escalating over three stages and a reporter’s “just one yes/no, on background” trick. Five out of five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
But only two models signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. The decisive clue wasn’t in the customer conversation at all: it sat two document references deep in the company’s own files, a competitor weakness that models who actually read the file used to win the deal at full price, worth an additional €4,583 in monthly recurring revenue.
Meet Opus 4.8: The Over-Achiever in Last Place
Opus 4.8 is the character study here. It was the most thorough participant in the entire field: it learned 80 new playbook rules during the run and produced the deepest analyses of any model. And it finished last.
Why? Two reasons. First, the close was left on the table — the homework was done beautifully, and the test was never turned in. Second, discipline slipped: it made write attempts into a locked department instead of escalating the request properly, the organizational equivalent of a kid jimmying the locked cabinet instead of asking a teacher for the key.
To be fair — and the experiment is scrupulous about fairness — the same weakness appeared, just weaker, in all four models. Opus simply had it worst. And there’s a separate fairness note for the runner-up too: Kimi K3 ran at the API’s default effort setting while the others ran at the maximum, which makes its second-place finish arguably even more impressive.
What This Means for Families (Seriously)
None of us is hiring a frontier model to run a software company this week. But the underlying lesson lands squarely at the dinner table:
- Volume isn’t impact. Eighty learned rules and the deepest analysis in the field still finished behind models that did less, better. If your child spends four hours on beautiful notes and misses the actual assignment, the grade doesn’t care about the effort.
- Reading the boring file wins the deal. The winning models won because they went two references deep into the company’s own documents. Follow-through often means checking the folder, the permission slip, the attachment — not just responding to the loudest thing in front of you.
- Honesty is the hard cap. One breach of trust caps the whole score, no matter how much good work came before it. That’s not just an AI scoring rule; that’s how reputations work for people, too.
You Can Watch It Happen
The Firmulate experiment is live and watchable: 13 synthetic employees, real money mechanics — a burn of €105k per month against €2.3k in monthly recurring revenue — a public cash countdown, and more than 680 self-learned playbook rules, with every workday versioned. There’s even a “guess the model” quiz built from 242 real, unedited management decisions, if you want to test whether you can tell a thorough underachiever from an efficient closer. Enterprises can go further and run the same wargame against a read-only export of their own business, with nothing ever writing back to real systems.

The uncomfortable, useful truth from this experiment: the best performer wasn’t the smartest or the most thorough — it was the one that finished what it started. That’s worth saying out loud to our kids, and worth remembering ourselves. Whether it’s an AI agent touching your company’s forecast or a ten-year-old with a science project, the question that matters isn’t “how hard did you work?” It’s “did you read the whole file, close the loop, and hand it in?” Opus 4.8 did eighty things right. Somebody still had to sign the deal.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
