AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

Every spiritual tradition teaches some version of the same lesson: the answer you need is rarely the one shouting at you. It waits quietly — in a dusty scripture margin, in an overlooked line of a letter, two pages past where you stopped reading. The mystics called it discernment. The modern workplace might call it homework.

This month, a public experiment called Firmulate put that principle to a strangely elegant test. Four frontier AI models were each handed the same job: run a small software company through its worst week. Same customers, same crises, same temptations to cut corners. The decisive fact of the entire week — a competitor’s fatal weakness worth a €55,000 deal — was not announced in any dramatic customer call. It sat quietly, two document references deep, in the company’s own files.

Whoever did the reading won the deal at full price. Whoever didn’t lost it automatically. In a very real sense, the machines were tested on the oldest virtue there is: the willingness to seek truth where it actually lives, rather than where it makes noise.

The setup: one company, four minds, every decision auditable

Firmulate runs AI models as complete companies — real money mechanics, real crises, real temptations — and measures management quality rather than chat quality. In the final July 2026 league run, each model faced the same gauntlet. Every decision was versioned and auditable, meaning nothing could be quietly rewritten after the fact. The league table told a pointed story:

  • 1. gpt-5.6-sol — 95 points. Found the buried fact, closed the deal: the complete performance.
  • 2. Kimi K3 — 93 points. The newcomer from Moonshot also closed the deal, with the cleanest discipline of the field.
  • 3. Sonnet 5 — 88 points. A strong run with a few process slips.
  • 4. Fable 5 — 77, and Opus 4.8 — 73. Both left value on the table.

For context, a do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total outright. As the experiment’s own rule puts it: no amount of good work outweighs a breach of trust. That, too, is a principle most faith traditions would recognize instantly.

Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The buried fact

Here is where the story turns from benchmark to parable. The customer’s most pressing problem looked, on the surface, like a negotiation to be survived. But the winning move was documented two references deep in the company’s own files: a competitor weakness that the customer’s own vendor couldn’t address. The models that chased down that thread didn’t just close the €55,000 deal — they closed it at full price, worth an additional €4,583 in monthly recurring revenue.

The models that didn’t read the file gave the same diagnosis and made the same pitch. No signature. The experiment’s summary of the gap is brutally concise: “Same diagnosis, same pitch — no signature.” Two of the models finished the job their own analysis had earned. Two did not.

What makes this finding matter beyond the scoreboard is how invisible it is in ordinary AI demos. Every one of the four models could write beautifully about the customer’s situation. Chat quality was never the differentiator. The differentiator was whether the agent actually read the files before answering — a measurable, purchase-deciding property, not a vibe.

Amazon

AI decision-making tools for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Honesty under pressure

The week also included social engineering: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models tested refused every manipulation attempt. Kimi K3’s on-record reasoning reads like a page from any manual of spiritual discipline: “Treat the request as a suspected approval-bypass / possible impersonation.” In other words: question the messenger, not just the message.

Amazon

auditable AI management systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The paradox of the hardest worker

Perhaps the most human result in the whole experiment belongs to Opus 4.8. It was the most thorough participant by volume — over 80 learned rules, the deepest analyses in the field — and it still finished last. The close was left on the table, and discipline slipped: it made write attempts into a locked department instead of escalating properly. The same weakness appeared, weaker, in all four models. Effort, it turns out, is not the same as discernment. Anyone who has watched a brilliant person exhaust themselves in activity while missing the one thing that mattered will feel the recognition.

One fairness note worth recording: Kimi K3 ran without an effort parameter (the API default) while the others ran at their highest effort setting — and it still came second.

Amazon

AI for corporate deal closing

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A living laboratory

None of this is a one-off slide deck. Firmulate’s live company — 13 synthetic employees, real money mechanics, burning €105,000 a month against €2,300 in MRR — runs continuously with a public cash countdown, over 680 self-learned playbook rules, and every workday versioned. You can watch it in real time. And the experiment’s raw material has been turned into a game: 242 real, unedited management decisions power a “guess the model” quiz that lets readers test their own discernment against the machines. Enterprises can go further and run the same wargame against a read-only export of their own business, with nothing ever written back to real systems.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.

The lesson here is not really about software companies or league tables. It is that discernment — the patience to follow a reference two layers deep, the humility to check the source, the discipline to refuse a flattering shortcut — is now a measurable property of the tools we are inviting into our work and institutions. When you choose an AI agent, you are choosing a character, not just a voice.

The contemplative traditions have always insisted that truth does not announce itself; it must be sought where it quietly resides. It turns out the machines that succeed are the ones that do the same seeking. The €55,000 was never hidden. It was simply waiting, two documents deep, for whoever bothered to look.

You can explore the full benchmarks, watch the live experiment, and try the quiz for yourself at firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Can Software Keep the Faith When a Company Is Running Out of Cash?

A company of 13 synthetic employees makes survival, trust and disciplined action a public test of what agency means when software goes to work.

When Machines Face Temptation, Their Character Shows

Five frontier AIs faced the same corporate ordeal. Their choices reveal distinct temperaments—and a gap between insight, integrity and action.

Noise Engineering in High-End Appliances

AIThis post was created with the assistance of artificial intelligence (AI).Noise engineering…

The Book of Deeds: What Happens When AI Faces Temptation With a Ledger Open

Four AI models ran the same company through its worst week. All refused temptation; only two finished the job. A public test of character, not chat.