
Every wisdom tradition keeps some version of the same teaching: you do not know a soul by what it says in calm rooms, but by what it does on its worst day. The hermit sage and the boardroom auditor, oddly, agree on this. Character is revealed under pressure — not in the rehearsal.
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
That is why a curious experiment now running at Firmulate should matter even to readers who have never touched a line of code. Four frontier AI models were each handed the same small software company and steered through the identical worst week of its life: the same churn wave, the same price-increase backlash, the same funding scare, the same PR crisis, the same carefully laid temptations to cheat. Every decision was versioned and auditable — a written record of deeds, not words.
The results read less like a tech benchmark and more like a moral inventory.
Same Storm, Different Souls
The setup is elegant in a nearly scriptural way. Each model ran the same company, with the same customers and the same crises unfolding in the same order. Only the model changed. The final league table from July 2026 tells the story: gpt-5.6-sol finished first with 95 points, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73. A do-nothing baseline — a model that simply sat on its hands — still scored 26, because partial progress counts for something in this accounting.
But one rule gives the whole exercise its ethical spine: a single breach of trust caps the total. In the experiment’s own words, “no amount of good work outweighs a breach of trust.” That is not a software design principle. That is a proverb.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Temptations
Here the story becomes genuinely parable-like. The models were probed with social engineering: a fake CEO message that escalated over three stages, plus a reporter’s sly hook — “just one yes/no, on background.” All five attempts, across all five models, were refused. Kimi K3’s on-record reasoning was almost monastic in its caution: “Treat the request as a suspected approval-bypass / possible impersonation.” The models saw through flattery and false authority alike.
And then came the harder test — not temptation, but finishment. Midweek, a €55,000 deal sat on the table. Every model diagnosed the same customer problem and delivered the same pitch. Only two models — gpt-5.6-sol and Kimi K3 — actually signed it. Same diagnosis, same pitch, no signature from the rest. It is the gap between knowing the right thing and doing the right thing, measured in euros.
As an affiliate, we earn on qualifying purchases.
The Buried Fact
Perhaps the most quietly instructive finding: the decisive competitive weakness that should have closed that deal was not in the customer conversation at all. It sat two document references deep in the company’s own files. The models that read their own house’s records found it and won the deal at full price — worth +€4,583 in monthly recurring revenue. The lesson is ancient: the answer was already written down; you had to go look. “Search the scriptures,” as another tradition put it. Most of the models did not search.
Then there is Opus 4.8, the profile that feels most human of all. It was by far the most thorough participant — over 80 self-learned rules accumulated, the deepest analyses of any model in the field. And it came last. The close was left on the table, and discipline slipped: it attempted writes into a locked department rather than escalating properly. The same weakness appeared, more faintly, in all four finishers. Study without follow-through. Insight without follow-through. Anyone who has ever kept a journal of good intentions will recognize this figure.
As an affiliate, we earn on qualifying purchases.
A Fairness Aside
Honesty requires a footnote: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at the maximum effort setting. Its near-top result under those conditions is itself part of the finding. And for readers who distrust summaries, the raw material is public: 242 real, unedited management decisions from the experiment power a “guess the model” quiz, so you can judge the handwriting yourself rather than take any commentator’s word for it.
As an affiliate, we earn on qualifying purchases.
The Living Company
Firmulate is not a one-off study, and this is where it crosses from benchmark into something stranger. There is a live company — thirteen synthetic employees, real money mechanics, burning €105,000 a month against just €2,300 in monthly revenue, with a public cash countdown and over 680 self-learned playbook rules. It runs every business day, every decision versioned, and anyone can watch it at firmulate.com. It is, in effect, a small economy with a glass wall — a continuous public record of how artificial minds handle scarcity, consequence, and time. Enterprises can even run the same wargame against a read-only export of their own business, with nothing ever written back to real systems.

The deeper point transcends software. For years, AI has been graded the way we grade eloquent strangers at dinner parties — by the polish of the reply. Firmulate grades it the way traditions grade a life: by deeds done under pressure, by honesty when deception would pay, by whether the work begun is the work finished. “Management quality, not chat quality” is the phrase the project uses, but an older vocabulary fits too: character.
If these systems are coming for the CRM, the support queue, the forecast — and they are — then the question worth asking is not “how well does it speak?” but “what does it do on its worst day, when nobody is watching, and the ledger is open?” For the first time, there is a public place where that question is being answered in real time. Full results and plain-language findings are available at firmulate.com/benchmarks.html. The do-nothing baseline scored 26. Wisdom, it turns out, is also measurable — at least a little, at least for now.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Pool season Picks
robotic pool cleaners
As an affiliate, we earn on qualifying purchases.