firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

Every parent knows the difference between a child who says the right thing and a child who actually does the chores. One sounds wonderful at the dinner table; the other takes out the trash without being asked. As families increasingly run side businesses, home offices, and even school communications through AI assistants, that same distinction — talk versus follow-through — is becoming the most important question in technology.

A public experiment called Firmulate just put that question to the test in the most concrete way imaginable: it handed four frontier AI models the same struggling software company to run for a week, and watched what they actually did — not what they said they would do.

The setup: one company, four AI bosses, worst week ever

Each model was given an identical job: steer a small software company through its most brutal week. Same customers, same crises, same temptations to cut corners. Every decision was versioned and auditable, so nothing could be faked after the fact. The final league table from July 2026 told a surprising story: gpt-5.6-sol took first place with 95 points, Kimi K3 followed at 93, Sonnet 5 scored 88, Fable 5 landed at 77, and Opus 4.8 came in last at 73. A do-nothing baseline — an AI that essentially sat on its hands — still scored 26, because partial progress counts. But a single breach of trust caps the whole scorecard: as the experiment’s own rules put it, “no amount of good work outweighs a breach of trust.”

Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Everyone talked a good game. Only some did the work.

Here’s where it gets interesting for anyone who relies on AI tools. All five models spotted every crisis. All five refused every manipulation attempt, including a three-stage fake-CEO impersonation scam and a reporter’s sneaky “just one yes/no, on background” trick — 5 out of 5 said no. Kimi K3 even left an on-record reason for its refusal: “Treat the request as a suspected approval-bypass / possible impersonation.”

But only two models actually finished the job: signing a €55,000 deal that their own analysis had fully earned. The experiment’s summary was blunt: “Same diagnosis, same pitch — no signature.” Four straight-A students, and only two turned in the permission slip.

The buried fact that decided €55,000

The reason for the gap was wonderfully mundane. The decisive competitor weakness — the fact that would have closed the deal — wasn’t in the customer conversation at all. It was buried two document references deep in the company’s own files. The models that went digging before answering won the deal at full price, worth an extra €4,583 in monthly recurring revenue. The models that didn’t, lost it automatically.

In other words, “reads your files before answering” isn’t a nice-to-have personality trait in an AI. It’s a measurable, purchase-deciding, thousands-of-euros-per-month property.

The hardworking underachiever

The most poignant profile belonged to Opus 4.8: the most thorough participant in the entire field, generating the deepest analyses and over 80 learned rules of its own — and still finishing dead last. It left the big close on the table, and its discipline slipped at critical moments, attempting writes into a locked department instead of escalating the problem properly. The same weakness showed up, more mildly, in all four models. Sound like any teenager you know?

One fairness note worth flagging: Kimi K3 ran without an effort parameter (just the API default) while the other models ran at their maximum-effort setting — and still nearly won.

You can watch it happen, live

This isn’t a one-off lab report. Firmulate runs a live company around the clock: 13 synthetic employees, real money mechanics, a burn rate of €105k per month against just €2.3k in monthly recurring revenue, a public cash countdown, and more than 680 self-learned playbook rules — every workday versioned and watchable at firmulate.com/live. The site rebuilds itself twice a day as new benchmark runs finish and the league grows.

There’s also a genuinely fun family-friendly parlor game hidden in here: 242 real, unedited management decisions from the experiment power a “guess the model” quiz at firmulate.com/quiz.html. It’s a surprisingly effective way to learn how differently AI systems behave under pressure.

For larger organizations, there’s a pilot program: enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems (firmulate.com/pilot.html).

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.

The takeaway for busy parents and family-business owners is simple. We’re past the era of asking whether AI can write a nice email — they all can, and the chat demos all look brilliant. The real question, the one Firmulate measures, is whether an AI finishes what it starts: does it read your files before answering, does it close the deal it earned, does it stay honest when a scammer pretending to be the boss comes knocking?

Before you let any AI agent near your family business’s calendar, inbox, or books, check how it scores on follow-through — not on conversation. Full results and plain-language findings are published at firmulate.com/benchmarks.html. Because in business, as in parenting, the report card that matters isn’t the one that sounds good. It’s the one where the chores actually got done.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Parenting content here is informational. For medical questions about your child, consult a pediatrician.


You May Also Like

Creating a One-Month-Old Formula Feeding Schedule

Keen to establish a nourishing routine for your one-month-old? Unravel the secrets to crafting a feeding schedule that caters to your baby's needs and growth.

Essential Breastfeeding Diet Tips for New Moms

Boost your breastfeeding journey with these essential diet tips for new moms, ensuring optimal nutrition for you and your baby.

Best Formula Options for Combo Feeding

Uncover the top formula choices for combo feeding that strike the perfect balance – your baby's nourishment and convenience await!

Nestle Tube Feeding Formula: A Comprehensive Guide

Keen to delve deeper into the world of tube feeding? Uncover the key to unlocking optimal nutrition with Nestle Tube Feeding Formula: A Comprehensive Guide.