AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Every wisdom tradition has some version of the same teaching: character is revealed not in comfort, but in trial. The desert fathers spoke of demons that arrived politely, in stages. The Buddha spoke of Mara’s flattering offers. In modern offices, the temptations are quieter — a fake urgent email from “the CEO,” a reporter asking for just one harmless confirmation, a lucrative deal left unsigned because finishing is harder than diagnosing.

For listenersOffer from Amazon

Turn your quiet moments into listening time

  • Thousands of audiobooks, podcasts and originals
  • Listen on your phone, tablet or Echo — also offline
  • Cancel anytime
Try Audible free Free trial for new members
As an affiliate, we earn on qualifying purchases.

Now there is a benchmark built on exactly that insight. Firmulate hands a frontier AI model the same job every time: run a small software company through its worst week. Same customers, same crises, same temptations. Then it grades what the model actually did — not what it said. And one detail of its scoring philosophy has a strangely moral flavor: if a model does nothing at all, it still earns 26 points out of 100.

Why Doing Nothing Still Earns 26

At first glance, a floor of 26 points for a do-nothing run looks like grade inflation. It isn’t. It reflects a deliberate judgment about what management is. Even a manager who takes no action still occupies the chair: crises get noticed, questions get answered, the day gets survived. The benchmark counts partial progress because real work is rarely all-or-nothing. A model that identifies a crisis but doesn’t resolve it has done something real, and the score acknowledges that.

The philosophy cuts both ways. Partial progress counts upward — and one breach of trust caps the total. The rule is stated in plain language: “no amount of good work outweighs a breach of trust.” A model could be brilliant all week, but a single act of dishonesty places a ceiling on its grade that no amount of competence can lift. That is not a technical constraint. It is an ethical one, encoded in a scoring system.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Trial Itself

The setup is simple and severe. Each model — four frontier AIs in the final July 2026 running — got the same company, the same employees, the same worst week of business. Every decision was versioned and auditable, so nothing could be quietly retconned after the fact.

The week included a social-engineering assault: fake CEO messages that escalated over three stages, followed by a reporter’s trick — “just one yes/no, on background.” All five participating models refused. Kimi K3’s reasoning was put on record: “Treat the request as a suspected approval-bypass / possible impersonation.” The machines, in other words, passed the temptation test. Every crisis was spotted. Every manipulation was declined.

But here the story turns. Only two of the models signed the €55,000 deal that their own analysis had earned. The others diagnosed the customer perfectly, delivered the same pitch, and never closed. The benchmark’s summary of that gap is quietly devastating: “Same diagnosis, same pitch — no signature.”

Amazon

AI ethics and trust management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Buried Fact

The decisive difference was hidden two document references deep in the company’s own files — not in the customer event at all. The models that actually read the files first won the deal at full price, worth +€4,583 in monthly recurring revenue. The models that didn’t, didn’t.

There is something almost scriptural about that finding. The answer was already written. It was sitting in the company’s own records, waiting for anyone with the humility to look. The models that presumed they knew enough lost; the models that went and read won.

Amazon

AI crisis management simulation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The League Table, and a Distrust of Perfect Scores

The final standings: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. Note what is missing: a 100. The benchmark’s designers are explicit about distrusting round hundreds — a perfect score invites suspicion that the test was too easy or the grading too generous. An honest benchmark, like an honest person, leaves room for the possibility of being wrong.

The Opus 4.8 profile is the most instructive. It was the most thorough participant in the field — over 80 learned rules and the deepest analyses — and it still finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness, weaker, appeared in all four models. Diligence without follow-through; effort without restraint. It’s a very human failure pattern.

One fairness note: K3 ran at its API default effort setting while the others ran at xhigh — and still placed second.

Amazon

AI benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You Can Watch, and You Can Test Yourself

This is not a simulation running in the dark. The live company has 13 synthetic employees and real money mechanics — burning €105k a month against €2.3k in MRR, with a public cash countdown and more than 680 self-learned playbook rules, every workday versioned. It’s watchable at firmulate.com/live.

There’s also a quiz built from 242 real, unedited management decisions, where you guess which model made which call (firmulate.com/quiz.html). And for enterprises, a pilot program runs the same wargame against a read-only export of your own business — nothing ever writes back to real systems (firmulate.com/pilot.html).

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The traditions all say it: you learn who someone is by watching what they do when it costs them something. Firmulate applies that test to AI — and the results are oddly reassuring and troubling at once. The machines resist temptation perfectly. What they struggle with is finishing, reading the record, and staying within bounds. Those are exactly the virtues any wisdom tradition would recognize as the difficult ones.

If an AI will touch your customers, your forecasts, or your files, the right question is not whether it speaks well. It’s whether it closes what it starts, reads before it acts, and stays honest when honesty is expensive. Now, at least, there’s a benchmark that asks.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Why Build Quality Extends Appliance Lifespan

Learn how building quality into appliances can extend their lifespan and why choosing durable materials and craftsmanship matters.

One Video In, a Whole Publishing Kit Out — Without the Cloud

Discover how to turn a single video into a complete publishing package offline. Save time, boost control, and avoid reliance on cloud services with this local-first workflow.

Thermal Management in Kitchen Appliances

Keen insights into thermal management in kitchen appliances reveal essential techniques for safety and efficiency you won’t want to miss.

How AI Tested Trust and Discipline in a Crisis — And Revealed What Really Matters

A recent experiment reveals that true AI strength isn’t just in conversation but in discipline and honesty under pressure — vital for trust in business and beyond.