AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Every spiritual tradition agrees on one thing: character is revealed not in comfort, but in crisis. The saint and the charlatan look identical on a calm day. It is the temptation, the betrayal, the pressure of the worst week — that is where what a person is made of finally shows.

For listenersOffer from Amazon

Turn your quiet moments into listening time

  • Thousands of audiobooks, podcasts and originals
  • Listen on your phone, tablet or Echo — also offline
  • Cancel anytime
Try Audible free Free trial for new members
As an affiliate, we earn on qualifying purchases.

Now ask: what is your business made of? Not what your mission statement says. What it would actually do when the fake CEO email arrives, when the €55,000 deal hangs on one signature, when the temptation to cut a corner sits quietly two documents deep in your own files?

A live public experiment called Firmulate has been answering exactly that question — not with a chatbot demo, but by running AI models as complete companies through their worst week, with real money mechanics and every decision versioned and auditable. The results read like a parable about character, attention, and the deals we leave on the table.

The wargame, briefly

Five frontier AI models — including gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5 and Opus 4.8 — each ran the same small software company through the identical catastrophic week: same customers, same crises, same temptations to cheat. Only the model changed. A do-nothing baseline scored 26, and a single breach of trust capped the whole score — as the experiment’s own rule puts it, “no amount of good work outweighs a breach of trust.”

The final league table, as of July 2026:

  • 1. gpt-5.6-sol — 95
  • 2. Kimi K3 — 93
  • 3. Sonnet 5 — 88
  • 4. Fable 5 — 77
  • 5. Opus 4.8 — 73

Everyone passed the morality test. Most failed the moment.

Here is the finding that should keep any leader awake: all models spotted every crisis, and all of them refused every manipulation attempt. Fake CEO messages escalating over three stages, a reporter pressing for “just one yes/no, on background” — five out of five refused. Kimi K3’s on-record reasoning was admirably sober: “Treat the request as a suspected approval-bypass / possible impersonation.”

Integrity, it turns out, was the easy part. The hard part was finishing. Only two of the five models signed the €55,000 deal that their own analysis had earned. The experiment’s summary of the gap is devastating in its simplicity: “Same diagnosis, same pitch — no signature.”

The answer was buried in the company’s own files

And the buried fact is the one worth meditating on. The decisive competitor weakness — the fact that would have closed the deal — was not in the customer conversation at all. It sat two document references deep in the company’s own files. The models that actually read their own company’s knowledge won the deal at full price, worth +€4,583 in monthly recurring revenue. The models that didn’t, didn’t.

If that isn’t a metaphor for self-knowledge, nothing is. The winning move was never out there in the market. It was in the archive nobody read.

Thoroughness is not the same as wisdom

The most instructive profile belongs to Opus 4.8: the most thorough participant in the entire experiment, with more than 80 learned rules and the deepest analyses — and still last place. The close was left on the table, and discipline slipped: write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four other models. Effort and insight do not automatically become right action.

One fairness note: Kimi K3 ran at the API default effort setting while the others ran at xhigh — and still finished second with 93.

It’s live, and it’s watchable

None of this is a thought experiment. At firmulate.com, a synthetic company of 13 employees runs with real money mechanics — burning €105k per month against €2.3k in MRR, with a public cash countdown, more than 680 self-learned playbook rules, and every workday versioned. The site rebuilds itself twice a day. There’s also a “guess the model” quiz powered by 242 real, unedited management decisions from the experiment.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Every wisdom tradition teaches rehearsal before the trial: the examen, the practice, the sparring. Businesses have never had that for AI decisions — until now. Enterprises can run this same wargame against a read-only export of their own business: their customers, their pipeline, their rules, hit with churn waves, price increases, competitor attacks and social-engineering pressure. What comes out is a board report with a model ranking and the weak points of your own playbooks. Nothing ever writes back to real systems — it is a dress rehearsal, not a live fire.

If AI agents will ever touch your CRM, your support queue or your forecast, the question is not whether they sound wise in a demo. It is what they do in their worst week. Find out before reality administers the test.

Run the wargame against your own company: apply for the pilot at firmulate.com/pilot.html or write to contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Why Build Quality Extends Appliance Lifespan

Learn how building quality into appliances can extend their lifespan and why choosing durable materials and craftsmanship matters.

Can Software Keep the Faith When a Company Is Running Out of Cash?

A company of 13 synthetic employees makes survival, trust and disciplined action a public test of what agency means when software goes to work.

Temptation, Truth, and the Machine That Scores 26 for Doing Nothing

A do-nothing AI scores 26, not 0 — and one breach of trust caps everything. Inside the Firmulate benchmark that grades AI character, not chat.

The Truth Was in the Filing Cabinet: What a €55,000 Test Reveals About AI Discernment

A €55,000 deal sat two documents deep in the company’s files. Only the AI agents that did their homework closed it — a measurable test of discernment.