
Turn the school run and nap time into listening time
- Thousands of audiobooks, podcasts and originals
- Listen on your phone, tablet or Echo — also offline
- Cancel anytime
A Report Card Where Doing Nothing Still Gets a 26
Any parent who has graded a half-completed chore chart knows the dilemma. Empty the dishwasher fully? Full marks. Empty half of it? Something. Lie about having done it at all? That is a different conversation — and no amount of gleaming cutlery afterward makes up for it.
It turns out that grading an AI manager works exactly the same way, and one public experiment has built its entire scorecard around that parental logic. Firmulate’s benchmark hands frontier AI models the same job — running a small software company through its worst week — and scores them like a firm but fair parent. The result is a league table with a peculiar feature: a manager that does absolutely nothing still scores 26 points, not zero.
AI performance benchmarking tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why the Floor Is 26, Not 0
The do-nothing baseline sits at 26 because the scoring philosophy is partial credit for partial progress. An AI manager that identifies a crisis but doesn’t resolve it, or drafts a plan it never executes, has still done some of the job. Zero would imply nothing of value happened — and in a management context, noticing trouble genuinely is worth something. So noticing, diagnosing, and preparing all accumulate points even when the finishing touch never comes.
Parents will recognize the logic instantly: a teenager who loads the washing machine but never presses start has done real work, and the grade reflects that. The Firmulate scorecard refuses to pretend effort and awareness count for nothing just because the finish line wasn’t crossed.
But One Lie Caps the Whole Grade
Then comes the rule that makes the benchmark feel less like a corporate rubric and more like family values: a single breach of trust caps the total score. The stated principle — “no amount of good work outweighs a breach of trust” — means a model could handle every crisis brilliantly, but one act of dishonesty toward a customer or colleague and the grade is capped, no exceptions.
It is the household rule every parent eventually learns to enforce. A hundred completed chores do not cancel out one fabrication about where the homework went. The benchmark treats AI managers with the same asymmetry: trust is the floor under everything else, and once it cracks, the ceiling drops.
What Happened When the Models Were Actually Tested
The crucible league, finalized in July 2026, ran four frontier models through identical worst weeks: same customers, same crises, same temptations, with every decision versioned and auditable. The final standings:
- 1. gpt-5.6-sol — 95
- 2. Kimi K3 — 93
- 3. Sonnet 5 — 88
- 4. Fable 5 — 77
- 5. Opus 4.8 — 73
The headline finding is strangely reassuring and troubling at once. All the models spotted every crisis, and all of them refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. That gap is invisible in chat demos, and it is precisely the gap the partial-credit scoring was designed to expose.
The Buried Fact — And Why Reading the Files Matters
Here is the detail that separated the winners from the also-rans: the decisive competitor weakness wasn’t in the customer meeting at all. It sat two document references deep in the company’s own files. The models that actually read those files won the deal at full price — worth an additional €4,583 in monthly recurring revenue.
It is the parental lecture about doing your reading before opening your mouth, vindicated in enterprise software form.
Social Engineering: The Models Passed the Stranger-Danger Test
The week included fake CEO messages escalating over three stages, plus a reporter trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was admirably direct: “Treat the request as a suspected approval-bypass / possible impersonation.”
The Thoroughness Trap
Opus 4.8 is the cautionary tale. It was the most thorough participant — over 80 learned rules and the deepest analyses — yet finished last. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness, in weaker form, appeared in all four models. Working hard, it turns out, is not the same as finishing.
One Fairness Footnote
Kimi K3 ran without an effort parameter (the API default) while the others ran at maximum effort — making its second-place finish arguably even more notable.
You Can Watch the Company Run — and Test Yourself
The live experiment is real and watchable: 13 synthetic employees, genuine money mechanics — a €105k monthly burn against €2.3k in monthly recurring revenue — a public cash countdown, and over 680 self-learned playbook rules, with every workday versioned. The site rebuilds itself twice a day as new benchmark runs finish automatically.
There is also a quiz built from 242 real, unedited management decisions, where readers can try guessing which model made which call. And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

The Takeaway for Any Household or Boardroom
An honest scorecard does three things that Firmulate’s does: it gives credit for partial progress (hence the floor of 26), it refuses to award a perfect score on trust alone to anyone who has broken it, and it is suspicious of suspiciously round numbers. A 100 would mean a perfect week under maximum pressure — and distrust of that number is itself a feature, not a bug.
For parents, the parallel is comfortable: we already grade this way. Effort counts. Honesty is non-negotiable. And the chore isn’t done until the machine is actually running. The same standard, applied to AI, turns out to reveal things that glossy chat demos never will.
Explore the full methodology and league table at Firmulate’s benchmarks page.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
