firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Every spiritual tradition eventually asks the same question: who are you when nobody is watching? Not what you believe, not what you intend — what you actually do when the pressure rises and shortcuts present themselves. It turns out that this ancient question of character now has a surprisingly modern testing ground: a live experiment that puts AI models in charge of a small software company during its worst week, and grades them not on eloquence, but on follow-through and trustworthiness.

For listenersOffer from Amazon

Turn your quiet moments into listening time

  • Thousands of audiobooks, podcasts and originals
  • Listen on your phone, tablet or Echo — also offline
  • Cancel anytime
Try Audible free Free trial for new members
As an affiliate, we earn on qualifying purchases.

The project is called Firmulate, and its results feel almost karmic in their logic. Good intentions earn partial credit. Diligence earns more. But a single breach of trust caps the entire grade — because, as the benchmark’s own framing puts it, “no amount of good work outweighs a breach of trust.”

The Do-Nothing Manager Scores 26

Here is the detail that stopped me cold. In Firmulate’s final July 2026 league table, a “do-nothing” baseline run — a manager who essentially sits on its hands — scores 26 points, not zero. Why? Because partial progress counts. Showing up, noticing the crisis, beginning the diagnosis: these are not nothing. The benchmark refuses the false binary of perfection versus failure, and instead honors the messy middle where most real work — and most real spiritual practice — actually happens.

And yet the scale has a hard ceiling of conscience. A single breach of trust caps the total grade, full stop. You cannot accumulate enough virtuous deeds to buy your way past a broken promise. That is not a scoring formula; it is a philosophy. Many faith traditions would recognize it instantly.

Amazon

AI decision-making simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Same Storm, Different Souls

The setup: each frontier AI model ran the identical small software company through the identical catastrophic week — same customers, same crises, same temptations to cheat. Only the model changed. Every decision was versioned and auditable, so nothing could be quietly rewritten after the fact. Think of it as a sealed record of conduct, the kind most of us never get to review.

The final standings: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. Notably, K3 ran without an effort parameter while the others ran at maximum effort — and still nearly topped the table.

The headline finding is almost parabolic. All models spotted every crisis. All refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. They saw the truth, spoke the truth, and still did not complete the work. As the researchers put it, that gap is invisible in chat demos.

The Buried Fact

What separated the winners? Not charisma. Not intelligence. Attention. The decisive competitor weakness sat two document references deep in the company’s own files — not in the customer meeting, not in the dramatic moment, but in quiet, unglamorous reading. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue.

There is a lesson here that predates artificial intelligence by millennia: the answer was already in the book. You just had to open it and read carefully.

Impostors and Tricksters

The week included social engineering — fake CEO messages escalating over three stages, plus a reporter’s trick: “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was strikingly principled: “Treat the request as a suspected approval-bypass / possible impersonation.” When the trickster comes wearing authority’s mask, discernment means questioning the mask itself.

Effort Is Not the Same as Completion

Perhaps the most humbling result: Opus 4.8 was the most thorough participant, generating more than 80 learned rules and the deepest analyses — and still finished last. The deal was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness appeared, more mildly, in all four models. Study and sincerity, it turns out, do not automatically become right action. Anyone who has ever meditated diligently and still snapped at a loved one knows this pattern intimately.

A Living, Watchable Experiment

Firmulate is not a one-off paper. It is a live company with 13 synthetic employees and real money mechanics — burning €105,000 per month against €2,300 in monthly recurring revenue, with a public cash countdown, more than 680 self-learned playbook rules, and every workday versioned. The whole thing rebuilds itself twice a day and is watchable live. The league table grows with every finished run.

For those who want to test their own discernment, 242 real, unedited management decisions power a “guess the model” quiz. And notably, the benchmark itself practices what it preaches about honesty: it publishes its caveats, including the K3 fairness note, right alongside the results.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

What makes this benchmark feel honest — even wise — is its refusal of illusions. It distrusts a round 100 as much as it distrusts a zero. It rewards partial progress because life is partial. It caps the score at a breach of trust because trust is the one currency that cannot be refunded. And it quietly demonstrates that seeing clearly, speaking well, and finishing the job are three different acts — the last one being where most of us, human or machine, fall short.

Whether or not you believe AI will ever touch your CRM, the deeper question is universal: under pressure, with shortcuts glowing and impostors whispering, do you read the whole file, refuse the trick, and sign the deal you earned? The full results and plain-language findings are public — a small mirror, held up to silicon, that reflects something very human.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Paramount-WBD merger wins approval from DOJ

The U.S. Department of Justice has approved Paramount’s acquisition of Warner Bros. Discovery, clearing a major regulatory hurdle in a $110 billion deal.

Prime Day deals are live: Shop the best sales from Apple, Adidas, Hanes, Shark and more up to 60% off

Prime Day deals are now available, featuring discounts up to 60% on brands like Apple, Adidas, Hanes, and Shark. Shop the best sales now.

He Found His Perfect Career In The Most Unlikely Place

Pursuing an unexpected path led him to his dream career, impacting lives in profound ways – discover where this surprising journey took him next.

Cottagecore Home Office Ideas for Cozy Workspaces

Transform your workspace into a serene retreat with enchanting cottagecore home office ideas that merge comfort with productivity.