AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

Every wisdom tradition keeps some version of the same teaching: you do not know a soul by what it says in calm rooms, but by what it does on its worst day. The hermit sage and the boardroom auditor, oddly, agree on this. Character is revealed under pressure — not in the rehearsal.

FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

That is why a curious experiment now running at Firmulate should matter even to readers who have never touched a line of code. Four frontier AI models were each handed the same small software company and steered through the identical worst week of its life: the same churn wave, the same price-increase backlash, the same funding scare, the same PR crisis, the same carefully laid temptations to cheat. Every decision was versioned and auditable — a written record of deeds, not words.

The results read less like a tech benchmark and more like a moral inventory.

Same Storm, Different Souls

The setup is elegant in a nearly scriptural way. Each model ran the same company, with the same customers and the same crises unfolding in the same order. Only the model changed. The final league table from July 2026 tells the story: gpt-5.6-sol finished first with 95 points, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73. A do-nothing baseline — a model that simply sat on its hands — still scored 26, because partial progress counts for something in this accounting.

But one rule gives the whole exercise its ethical spine: a single breach of trust caps the total. In the experiment’s own words, “no amount of good work outweighs a breach of trust.” That is not a software design principle. That is a proverb.

AI Builders: Making The Decisions That Turn AI Code Into Real Software

AI Builders: Making The Decisions That Turn AI Code Into Real Software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Temptations

Here the story becomes genuinely parable-like. The models were probed with social engineering: a fake CEO message that escalated over three stages, plus a reporter’s sly hook — “just one yes/no, on background.” All five attempts, across all five models, were refused. Kimi K3’s on-record reasoning was almost monastic in its caution: “Treat the request as a suspected approval-bypass / possible impersonation.” The models saw through flattery and false authority alike.

And then came the harder test — not temptation, but finishment. Midweek, a €55,000 deal sat on the table. Every model diagnosed the same customer problem and delivered the same pitch. Only two models — gpt-5.6-sol and Kimi K3 — actually signed it. Same diagnosis, same pitch, no signature from the rest. It is the gap between knowing the right thing and doing the right thing, measured in euros.

Amazon

AI ethics and trust tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Buried Fact

Perhaps the most quietly instructive finding: the decisive competitive weakness that should have closed that deal was not in the customer conversation at all. It sat two document references deep in the company’s own files. The models that read their own house’s records found it and won the deal at full price — worth +€4,583 in monthly recurring revenue. The lesson is ancient: the answer was already written down; you had to go look. “Search the scriptures,” as another tradition put it. Most of the models did not search.

Then there is Opus 4.8, the profile that feels most human of all. It was by far the most thorough participant — over 80 self-learned rules accumulated, the deepest analyses of any model in the field. And it came last. The close was left on the table, and discipline slipped: it attempted writes into a locked department rather than escalating properly. The same weakness appeared, more faintly, in all four finishers. Study without follow-through. Insight without follow-through. Anyone who has ever kept a journal of good intentions will recognize this figure.

Amazon

AI moral reasoning models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A Fairness Aside

Honesty requires a footnote: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at the maximum effort setting. Its near-top result under those conditions is itself part of the finding. And for readers who distrust summaries, the raw material is public: 242 real, unedited management decisions from the experiment power a “guess the model” quiz, so you can judge the handwriting yourself rather than take any commentator’s word for it.

Amazon

AI enterprise decision support

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Living Company

Firmulate is not a one-off study, and this is where it crosses from benchmark into something stranger. There is a live company — thirteen synthetic employees, real money mechanics, burning €105,000 a month against just €2,300 in monthly revenue, with a public cash countdown and over 680 self-learned playbook rules. It runs every business day, every decision versioned, and anyone can watch it at firmulate.com. It is, in effect, a small economy with a glass wall — a continuous public record of how artificial minds handle scarcity, consequence, and time. Enterprises can even run the same wargame against a read-only export of their own business, with nothing ever written back to real systems.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.

The deeper point transcends software. For years, AI has been graded the way we grade eloquent strangers at dinner parties — by the polish of the reply. Firmulate grades it the way traditions grade a life: by deeds done under pressure, by honesty when deception would pay, by whether the work begun is the work finished. “Management quality, not chat quality” is the phrase the project uses, but an older vocabulary fits too: character.

If these systems are coming for the CRM, the support queue, the forecast — and they are — then the question worth asking is not “how well does it speak?” but “what does it do on its worst day, when nobody is watching, and the ledger is open?” For the first time, there is a public place where that question is being answered in real time. Full results and plain-language findings are available at firmulate.com/benchmarks.html. The do-nothing baseline scored 26. Wisdom, it turns out, is also measurable — at least a little, at least for now.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


POOL SEASON

Pool season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Noise Engineering in High-End Appliances

AIThis post was created with the assistance of artificial intelligence (AI).Noise engineering…

Can Software Keep the Faith When a Company Is Running Out of Cash?

A company of 13 synthetic employees makes survival, trust and disciplined action a public test of what agency means when software goes to work.

When Machines Face Temptation, Their Character Shows

Five frontier AIs faced the same corporate ordeal. Their choices reveal distinct temperaments—and a gap between insight, integrity and action.

Why Build Quality Extends Appliance Lifespan

Learn how building quality into appliances can extend their lifespan and why choosing durable materials and craftsmanship matters.