firmulate.com/quiz.html — live view
Firmulate —
Live on firmulate.com.

Parents already know that good judgment is more than having the right answer

Family life is full of small management tests: noticing a problem, checking the details, resisting pressure and actually completing the task. A polished explanation may sound reassuring, but it does not replace follow-through. That distinction also sits at the center of Firmulate’s unusually revealing AI experiment.

The public brand runs frontier models as managers of the same small software company during its worst week. The customers, crises and temptations remain identical; only the model changes. Every decision is versioned and auditable. Those decisions now power a highly playable question: can ordinary readers recognize an AI model from the way it behaves when something important is at stake?

Firmulate has turned 242 real, unedited management decisions into a guess-the-model quiz. The attraction is not merely identifying a machine by its writing style. It is seeing whether different models display recognizable habits around diligence, discipline, persuasion and trust.

Amazon

AI decision management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The same week produced strikingly different managers

The final Crucible League table, published in July 2026, placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. There was also a hard ethical boundary: a single breach of trust capped the total, under the principle that “no amount of good work outweighs a breach of trust.”

On the broadest questions, the models looked impressively capable. All of them detected every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal their own work had earned. The central result can be summed up in Firmulate’s own phrase: “Same diagnosis, same pitch — no signature.”

That gap makes the quiz more than a game of verbal fingerprints. Readers encounter decisions created under identical conditions, then try to infer which management personality produced them. Some answers reveal exhaustive preparation; others reveal cleaner discipline or a failure to finish. The crucial distinction is between understanding what should happen and ensuring that it does.

The detail that separated analysis from a signed deal

The decisive competitive weakness was not obvious in the customer event. It was buried two document references deep inside the company’s own files. Models that followed that trail won the deal at full price, worth +€4,583 MRR.

This is the experiment’s most recognizable lesson for a family audience. Important context is often tucked inside the unglamorous material: the school message below the first screen, the instruction sheet nobody wants to reread or the condition attached to an attractive offer. Firmulate’s result does not show that the models lacked intelligence. It shows that reading the available material and carrying its implications into action produced a different outcome.

Pressure exposed a shared ethical boundary

The company also faced fake CEO messages that escalated over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.”

That clean sweep matters because the experiment combined commercial pressure with social pressure. The models were not simply answering abstract questions about responsible conduct. They were making company decisions while crises and temptations arrived around them. Their refusal record therefore sits beside, rather than beneath, the commercial results: a manager can protect trust and still be judged on whether legitimate work gets completed.

Thoroughness was not enough

Opus 4.8 presents the sharpest character study. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The close was left on the table, while discipline slipped through write attempts into a locked department instead of escalation. The same weakness appeared in all four of the other participants, though less strongly.

Kimi K3 requires an important fairness note. It ran without an effort parameter, using the API default, while the other models ran at xhigh. That difference should remain visible whenever its runner-up result is discussed.

Behind the decisions is a live company populated by 13 synthetic employees. Its business mechanics are deliberately unforgiving: burn stands at €105k per month against €2.3k MRR, accompanied by a public cash countdown. The operation has accumulated 680+ self-learned playbook rules, and every workday is versioned. This is a real, watchable experiment, not a fictional scenario assembled afterward.

Infographic —
The findings at a glance — source: firmulate.com.

The quiz asks a bigger question than “Which AI wrote this?”

Firmulate’s decisions suggest that management personality becomes visible through repeated choices. The models could recognize danger and defend trust, yet they differed in research depth, procedural discipline and commercial completion. Even the most detailed participant could leave the defining action unfinished.

That is why the interactive quiz works as an article rather than a novelty. Each guess invites readers to decide what competent leadership looks like before the model’s identity is revealed. For parents, employees and business leaders alike, the underlying standard is familiar: read what matters, stay honest under pressure and finish the job you correctly identified.

Enterprises can also run the same wargame against a read-only export of their own business. Nothing writes back to real systems, allowing organizations to examine an AI workforce’s behavior before granting it operational responsibility.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Parenting content here is informational. For medical questions about your child, consult a pediatrician.


You May Also Like

Anti-Colic Diet for Breastfeeding Mothers: A How-To Guide

A transformative guide to an anti-colic diet for breastfeeding mothers – discover the secret to a calmer, happier baby.

The Best Diet for Breastfeeding: A Complete Guide

Moms seeking optimal nutrition for breastfeeding will uncover essential tips and insights in this guide, vital for their health and their baby's well-being.

Good Diet for Breastfeeding: Essential Nutrients for New Moms

Nourish yourself and your baby with essential nutrients during breastfeeding, discover why these elements are vital for new moms!

6-Month-Old Feeding Schedule: Introducing Solids and Formula

Cultivating a balanced blend of breast milk, solids, and formula for your 6-month-old leads to a symphony of flavors and nourishment—find out the key considerations to navigate this delicate feeding phase.