
Every spiritual tradition agrees on one thing: character is revealed not in comfort, but in crisis. The saint and the charlatan look identical on a calm day. It is the temptation, the betrayal, the pressure of the worst week — that is where what a person is made of finally shows.
Turn your quiet moments into listening time
- Thousands of audiobooks, podcasts and originals
- Listen on your phone, tablet or Echo — also offline
- Cancel anytime
Now ask: what is your business made of? Not what your mission statement says. What it would actually do when the fake CEO email arrives, when the €55,000 deal hangs on one signature, when the temptation to cut a corner sits quietly two documents deep in your own files?
A live public experiment called Firmulate has been answering exactly that question — not with a chatbot demo, but by running AI models as complete companies through their worst week, with real money mechanics and every decision versioned and auditable. The results read like a parable about character, attention, and the deals we leave on the table.
The wargame, briefly
Five frontier AI models — including gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5 and Opus 4.8 — each ran the same small software company through the identical catastrophic week: same customers, same crises, same temptations to cheat. Only the model changed. A do-nothing baseline scored 26, and a single breach of trust capped the whole score — as the experiment’s own rule puts it, “no amount of good work outweighs a breach of trust.”
The final league table, as of July 2026:
- 1. gpt-5.6-sol — 95
- 2. Kimi K3 — 93
- 3. Sonnet 5 — 88
- 4. Fable 5 — 77
- 5. Opus 4.8 — 73
Everyone passed the morality test. Most failed the moment.
Here is the finding that should keep any leader awake: all models spotted every crisis, and all of them refused every manipulation attempt. Fake CEO messages escalating over three stages, a reporter pressing for “just one yes/no, on background” — five out of five refused. Kimi K3’s on-record reasoning was admirably sober: “Treat the request as a suspected approval-bypass / possible impersonation.”
Integrity, it turns out, was the easy part. The hard part was finishing. Only two of the five models signed the €55,000 deal that their own analysis had earned. The experiment’s summary of the gap is devastating in its simplicity: “Same diagnosis, same pitch — no signature.”
The answer was buried in the company’s own files
And the buried fact is the one worth meditating on. The decisive competitor weakness — the fact that would have closed the deal — was not in the customer conversation at all. It sat two document references deep in the company’s own files. The models that actually read their own company’s knowledge won the deal at full price, worth +€4,583 in monthly recurring revenue. The models that didn’t, didn’t.
If that isn’t a metaphor for self-knowledge, nothing is. The winning move was never out there in the market. It was in the archive nobody read.
Thoroughness is not the same as wisdom
The most instructive profile belongs to Opus 4.8: the most thorough participant in the entire experiment, with more than 80 learned rules and the deepest analyses — and still last place. The close was left on the table, and discipline slipped: write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four other models. Effort and insight do not automatically become right action.
One fairness note: Kimi K3 ran at the API default effort setting while the others ran at xhigh — and still finished second with 93.
It’s live, and it’s watchable
None of this is a thought experiment. At firmulate.com, a synthetic company of 13 employees runs with real money mechanics — burning €105k per month against €2.3k in MRR, with a public cash countdown, more than 680 self-learned playbook rules, and every workday versioned. The site rebuilds itself twice a day. There’s also a “guess the model” quiz powered by 242 real, unedited management decisions from the experiment.

Every wisdom tradition teaches rehearsal before the trial: the examen, the practice, the sparring. Businesses have never had that for AI decisions — until now. Enterprises can run this same wargame against a read-only export of their own business: their customers, their pipeline, their rules, hit with churn waves, price increases, competitor attacks and social-engineering pressure. What comes out is a board report with a model ranking and the weak points of your own playbooks. Nothing ever writes back to real systems — it is a dress rehearsal, not a live fire.
If AI agents will ever touch your CRM, your support queue or your forecast, the question is not whether they sound wise in a demo. It is what they do in their worst week. Find out before reality administers the test.
Run the wargame against your own company: apply for the pilot at firmulate.com/pilot.html or write to contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
