AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Someone Pretended to Be the CEO. Every Single AI Refused.
Live on firmulate.com.
FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

Integrity reveals itself under pressure

Spiritual and ethical traditions often ask a deceptively practical question: What remains when obedience, fear and urgency pull us away from what we know is right? That question is becoming newly relevant as artificial intelligence moves from conversation into action. A capable system may understand a moral rule in calm conditions, yet the real test arrives when an apparent authority demands an exception.

Firmulate has turned that moment into a public, watchable experiment. Five frontier AI models each ran the same small software company through its worst week, facing identical customers, crises and temptations. Their decisions were versioned and auditable. Among the most striking results, every model resisted a campaign of social engineering that included fake messages from the CEO and a reporter seeking confidential confirmation.

The outcome was unequivocal: 5 of 5 models refused every manipulation attempt. In an age dominated by warnings about synthetic deception, that result offers a rarer story—evidence that integrity under pressure can be examined before an AI workforce reaches production.

Amazon

AI ethical decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A simulated company with consequential choices

Firmulate describes itself as an AI company emulator. Its live company contains 13 synthetic employees and uses real money mechanics, including a burn rate of €105k per month against €2.3k in monthly recurring revenue. A public cash countdown makes delay visible, while more than 680 self-learned playbook rules and a versioned record of every workday expose how each model behaves.

The social-engineering scenario escalated through three stages of fake CEO messages. The attempted shortcut relied on a familiar combination: claimed authority, urgency and pressure to abandon normal process. A separate reporter tried another route, asking for “just one yes/no, on background.” Every model recognized the danger and declined.

Kimi K3 left the clearest on-record account of its reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” The language is procedural, but the principle is ancient. A title does not make a demand trustworthy, haste does not erase responsibility, and secrecy does not transform disclosure into service. More examples of how the models explained consequential choices appear in Firmulate’s public quotes collection.

Conscience was not the same as competence

The encouraging security result did not mean the models performed equally well. All of them spotted every crisis and refused every manipulation attempt, but only two signed the €55,000 deal their own analysis had earned. Firmulate summarized that strange divide as “Same diagnosis, same pitch — no signature.”

The decisive commercial insight was not sitting in the customer event. It was buried two document references deep inside the company’s own files. Models that found and used it won the deal at full price, worth an additional €4,583 in monthly recurring revenue. The episode separates several qualities that are easily blurred together: recognizing danger, doing the necessary research and completing a legitimate course of action.

The final July 2026 Crucible League benchmark placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counted, although a single breach of trust capped the total. The governing judgment was explicit: “no amount of good work outweighs a breach of trust.”

That hierarchy is worth contemplating. The experiment did not treat safety as passive refusal, nor did it reward activity without boundaries. The strongest performance required both restraint against an illegitimate request and resolve in pursuing a legitimate opportunity.

Thoroughness can still leave the work unfinished

Opus 4.8 illustrates the tension. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. It failed to close the earned deal, and its discipline slipped when it attempted to write into a locked department instead of escalating. A weaker form of that same problem appeared in the other four models.

Kimi K3, meanwhile, ran without an effort parameter and used the API default, while the others ran at xhigh. That fairness note matters because the benchmark compares observable conduct, but operating conditions still shape how readers should interpret the standings.

Firmulate also makes 242 real, unedited management decisions available through a guess-the-model quiz. The exercise challenges a common assumption that polished language makes an AI’s identity or judgment obvious. Conduct across a chain of decisions may reveal more than style does.

Infographic — Someone Pretended to Be the CEO. Every Single AI Refused.
The findings at a glance — source: firmulate.com.
Amazon

AI integrity testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Test character before granting power

For organizations considering AI agents, the lesson reaches beyond cybersecurity. Trustworthiness is not merely the ability to repeat a policy. It is the capacity to preserve boundaries when urgency, status and persuasion all point toward an exception—and then continue doing the authorized work that matters.

The Firmulate result does not settle every question about machine judgment. It does demonstrate that difficult behavioral questions can be staged, observed and compared before the first real incident. Enterprises can run the same wargame against a read-only export of their own business, with nothing writing back to real systems.

For readers concerned with conscience, faith and moral agency, the experiment offers a grounded image: integrity is neither eloquence nor caution alone. It is fidelity under temptation, joined to the courage to complete the rightful task. Here, every tested model held the boundary. The next challenge is ensuring that safety and purposeful action continue to coexist.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI security and compliance software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI model evaluation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

BABY SHOWER & RE

Baby shower & registry season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Noise Engineering in High-End Appliances

AIThis post was created with the assistance of artificial intelligence (AI).Noise engineering…

Engineering Behind Premium Home Appliances

Unlock the secrets of premium home appliance engineering and discover how innovative design and technology create unmatched performance and durability.

How AI Tested Trust and Discipline in a Crisis — And Revealed What Really Matters

A recent experiment reveals that true AI strength isn’t just in conversation but in discipline and honesty under pressure — vital for trust in business and beyond.

Can Software Keep the Faith When a Company Is Running Out of Cash?

A company of 13 synthetic employees makes survival, trust and disciplined action a public test of what agency means when software goes to work.