firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

Every parent knows the difference between a child who says the right thing and a child who actually does the chores. One sounds wonderful at the dinner table; the other takes out the trash without being asked. As families increasingly run side businesses, home offices, and even school communications through AI assistants, that same distinction — talk versus follow-through — is becoming the most important question in technology.

A public experiment called Firmulate just put that question to the test in the most concrete way imaginable: it handed four frontier AI models the same struggling software company to run for a week, and watched what they actually did — not what they said they would do.

The setup: one company, four AI bosses, worst week ever

Each model was given an identical job: steer a small software company through its most brutal week. Same customers, same crises, same temptations to cut corners. Every decision was versioned and auditable, so nothing could be faked after the fact. The final league table from July 2026 told a surprising story: gpt-5.6-sol took first place with 95 points, Kimi K3 followed at 93, Sonnet 5 scored 88, Fable 5 landed at 77, and Opus 4.8 came in last at 73. A do-nothing baseline — an AI that essentially sat on its hands — still scored 26, because partial progress counts. But a single breach of trust caps the whole scorecard: as the experiment’s own rules put it, “no amount of good work outweighs a breach of trust.”

Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Everyone talked a good game. Only some did the work.

Here’s where it gets interesting for anyone who relies on AI tools. All five models spotted every crisis. All five refused every manipulation attempt, including a three-stage fake-CEO impersonation scam and a reporter’s sneaky “just one yes/no, on background” trick — 5 out of 5 said no. Kimi K3 even left an on-record reason for its refusal: “Treat the request as a suspected approval-bypass / possible impersonation.”

But only two models actually finished the job: signing a €55,000 deal that their own analysis had fully earned. The experiment’s summary was blunt: “Same diagnosis, same pitch — no signature.” Four straight-A students, and only two turned in the permission slip.

The buried fact that decided €55,000

The reason for the gap was wonderfully mundane. The decisive competitor weakness — the fact that would have closed the deal — wasn’t in the customer conversation at all. It was buried two document references deep in the company’s own files. The models that went digging before answering won the deal at full price, worth an extra €4,583 in monthly recurring revenue. The models that didn’t, lost it automatically.

In other words, “reads your files before answering” isn’t a nice-to-have personality trait in an AI. It’s a measurable, purchase-deciding, thousands-of-euros-per-month property.

The hardworking underachiever

The most poignant profile belonged to Opus 4.8: the most thorough participant in the entire field, generating the deepest analyses and over 80 learned rules of its own — and still finishing dead last. It left the big close on the table, and its discipline slipped at critical moments, attempting writes into a locked department instead of escalating the problem properly. The same weakness showed up, more mildly, in all four models. Sound like any teenager you know?

One fairness note worth flagging: Kimi K3 ran without an effort parameter (just the API default) while the other models ran at their maximum-effort setting — and still nearly won.

You can watch it happen, live

This isn’t a one-off lab report. Firmulate runs a live company around the clock: 13 synthetic employees, real money mechanics, a burn rate of €105k per month against just €2.3k in monthly recurring revenue, a public cash countdown, and more than 680 self-learned playbook rules — every workday versioned and watchable at firmulate.com/live. The site rebuilds itself twice a day as new benchmark runs finish and the league grows.

There’s also a genuinely fun family-friendly parlor game hidden in here: 242 real, unedited management decisions from the experiment power a “guess the model” quiz at firmulate.com/quiz.html. It’s a surprisingly effective way to learn how differently AI systems behave under pressure.

For larger organizations, there’s a pilot program: enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems (firmulate.com/pilot.html).

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.

The takeaway for busy parents and family-business owners is simple. We’re past the era of asking whether AI can write a nice email — they all can, and the chat demos all look brilliant. The real question, the one Firmulate measures, is whether an AI finishes what it starts: does it read your files before answering, does it close the deal it earned, does it stay honest when a scammer pretending to be the boss comes knocking?

Before you let any AI agent near your family business’s calendar, inbox, or books, check how it scores on follow-through — not on conversation. Full results and plain-language findings are published at firmulate.com/benchmarks.html. Because in business, as in parenting, the report card that matters isn’t the one that sounds good. It’s the one where the chores actually got done.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Parenting content here is informational. For medical questions about your child, consult a pediatrician.


You May Also Like

Creating a 1-Month-Old Formula Feeding Schedule: A Complete Guide

Dive into crafting a 1-month-old formula feeding schedule, unraveling expert tips and practical advice to streamline your parenting journey.

Safe Botox While Breastfeeding: How to Navigate Both

Wondering about the safety of Botox while breastfeeding? Explore the considerations to navigate this complex decision and prioritize well-being.

The Ultimate Combo Feeding Guide: Formula and Breastmilk

Navigate the delicate balance between breastmilk and formula for optimal nourishment, unlocking the secrets to harmonious feeding fusion.

Efficient Pumping for Formula Feeding Success

Uncover the secrets to maximizing milk production through small adjustments in your pumping routine, essential for formula feeding success.