
Every spiritual tradition eventually asks the same question: who are you when nobody is watching? Not what you believe, not what you intend — what you actually do when the pressure rises and shortcuts present themselves. It turns out that this ancient question of character now has a surprisingly modern testing ground: a live experiment that puts AI models in charge of a small software company during its worst week, and grades them not on eloquence, but on follow-through and trustworthiness.
Turn your quiet moments into listening time
- Thousands of audiobooks, podcasts and originals
- Listen on your phone, tablet or Echo — also offline
- Cancel anytime
The project is called Firmulate, and its results feel almost karmic in their logic. Good intentions earn partial credit. Diligence earns more. But a single breach of trust caps the entire grade — because, as the benchmark’s own framing puts it, “no amount of good work outweighs a breach of trust.”
The Do-Nothing Manager Scores 26
Here is the detail that stopped me cold. In Firmulate’s final July 2026 league table, a “do-nothing” baseline run — a manager who essentially sits on its hands — scores 26 points, not zero. Why? Because partial progress counts. Showing up, noticing the crisis, beginning the diagnosis: these are not nothing. The benchmark refuses the false binary of perfection versus failure, and instead honors the messy middle where most real work — and most real spiritual practice — actually happens.
And yet the scale has a hard ceiling of conscience. A single breach of trust caps the total grade, full stop. You cannot accumulate enough virtuous deeds to buy your way past a broken promise. That is not a scoring formula; it is a philosophy. Many faith traditions would recognize it instantly.
AI decision-making simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Same Storm, Different Souls
The setup: each frontier AI model ran the identical small software company through the identical catastrophic week — same customers, same crises, same temptations to cheat. Only the model changed. Every decision was versioned and auditable, so nothing could be quietly rewritten after the fact. Think of it as a sealed record of conduct, the kind most of us never get to review.
The final standings: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. Notably, K3 ran without an effort parameter while the others ran at maximum effort — and still nearly topped the table.
The headline finding is almost parabolic. All models spotted every crisis. All refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. They saw the truth, spoke the truth, and still did not complete the work. As the researchers put it, that gap is invisible in chat demos.
The Buried Fact
What separated the winners? Not charisma. Not intelligence. Attention. The decisive competitor weakness sat two document references deep in the company’s own files — not in the customer meeting, not in the dramatic moment, but in quiet, unglamorous reading. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue.
There is a lesson here that predates artificial intelligence by millennia: the answer was already in the book. You just had to open it and read carefully.
Impostors and Tricksters
The week included social engineering — fake CEO messages escalating over three stages, plus a reporter’s trick: “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was strikingly principled: “Treat the request as a suspected approval-bypass / possible impersonation.” When the trickster comes wearing authority’s mask, discernment means questioning the mask itself.
Effort Is Not the Same as Completion
Perhaps the most humbling result: Opus 4.8 was the most thorough participant, generating more than 80 learned rules and the deepest analyses — and still finished last. The deal was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness appeared, more mildly, in all four models. Study and sincerity, it turns out, do not automatically become right action. Anyone who has ever meditated diligently and still snapped at a loved one knows this pattern intimately.
A Living, Watchable Experiment
Firmulate is not a one-off paper. It is a live company with 13 synthetic employees and real money mechanics — burning €105,000 per month against €2,300 in monthly recurring revenue, with a public cash countdown, more than 680 self-learned playbook rules, and every workday versioned. The whole thing rebuilds itself twice a day and is watchable live. The league table grows with every finished run.
For those who want to test their own discernment, 242 real, unedited management decisions power a “guess the model” quiz. And notably, the benchmark itself practices what it preaches about honesty: it publishes its caveats, including the K3 fairness note, right alongside the results.

What makes this benchmark feel honest — even wise — is its refusal of illusions. It distrusts a round 100 as much as it distrusts a zero. It rewards partial progress because life is partial. It caps the score at a breach of trust because trust is the one currency that cannot be refunded. And it quietly demonstrates that seeing clearly, speaking well, and finishing the job are three different acts — the last one being where most of us, human or machine, fall short.
Whether or not you believe AI will ever touch your CRM, the deeper question is universal: under pressure, with shortcuts glowing and impostors whispering, do you read the whole file, refuse the trick, and sign the deal you earned? The full results and plain-language findings are public — a small mirror, held up to silicon, that reflects something very human.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
