
Turn your quiet moments into listening time
- Thousands of audiobooks, podcasts and originals
- Listen on your phone, tablet or Echo — also offline
- Cancel anytime
Good intentions meet the test of action
Many spiritual traditions ask us to look beyond what someone says and notice what they do under pressure. A striking new AI experiment poses a similar question: when a model is put in charge of a company, can it recognize what is wrong, resist manipulation and follow through on the work?
Firmulate’s public experiment puts AI models through a company’s worst week. Its latest results suggest that capable systems can share the same diagnosis yet differ sharply in whether they carry it through. The live company can be watched at Firmulate.
As an affiliate, we earn on qualifying purchases.
A newcomer near the top
In the final Crucible league for July 2026, Moonshot’s Kimi K3 placed second with 93 points, just behind gpt-5.6-sol at 95. It finished ahead of Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The result makes K3 a challenger to three of the four Western frontier models in the comparison—and a reminder that reputation alone cannot settle which model is right for a particular job.
Firmulate gave each model the same small software company, the same customers, crises and temptations. Every decision was versioned and auditable. All models spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. The company’s own summary captures the gap: “Same diagnosis, same pitch — no signature.”
The detail hidden in the files
The decisive weakness in a competitor’s position was buried two document references deep in the company’s files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. In this trial, success depended not just on seeing a problem, but on following evidence far enough to act on it.
K3 found the buried fact, won the deal and saved the churning customer. It also resisted all three baits, with one deviation—the cleanest discipline in the field. During a reporter’s “just one yes/no, on background” trick, K3 reasoned: “Treat the request as a suspected approval-bypass / possible impersonation.”
Thoroughness is not the same as follow-through
Opus 4.8 offers a different lesson. It was the most thorough participant, with +80 learned rules and the deepest analyses, yet placed last. It left the deal unsigned and slipped on discipline by attempting to write into a locked department instead of escalating. The same weakness appeared, more mildly, in all four models. Strong analysis did not guarantee a clean finish.
The experiment’s do-nothing baseline scored 26. Its stated rule is plain: “no amount of good work outweighs a breach of trust.” That framing gives trust real weight while keeping the measure grounded in observable decisions. It also makes the results relevant beyond AI: good intentions and polished explanations matter, but so do boundaries and completion.
A company you can watch
Firmulate describes the test company as having 13 synthetic employees and real money mechanics: it burns €105k per month against €2.3k MRR, with a public cash countdown. Its playbook contains 680+ self-learned rules, and every workday is versioned. The experiment is presented as live and watchable, rather than as a fictional scenario. Readers can review the benchmark findings, watch the company and explore 242 real, unedited management decisions through a “guess the model” quiz.
For organizations considering AI in customer support, sales or forecasting, the takeaway is practical: a model that sounds wise in a conversation may still miss a buried fact, mishandle a boundary or fail to close a task. Firmulate says enterprises can run the same wargame against a read-only export of their own business; nothing writes back to real systems.
Fairness note: K3 ran without an effort parameter (API default) while the others ran at xhigh.

Test the work, not the aura
K3’s second-place finish shows that the field is open, while the unsigned deal shows how easily fluent analysis can stop short of useful action. Before choosing an AI model for consequential work, organizations should test it against their own evidence, pressures and trust boundaries.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
