firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

For listenersOffer from Amazon

Turn your quiet moments into listening time

  • Thousands of audiobooks, podcasts and originals
  • Listen on your phone, tablet or Echo — also offline
  • Cancel anytime
Try Audible free Free trial for new members
As an affiliate, we earn on qualifying purchases.

Good intentions meet the test of action

Many spiritual traditions ask us to look beyond what someone says and notice what they do under pressure. A striking new AI experiment poses a similar question: when a model is put in charge of a company, can it recognize what is wrong, resist manipulation and follow through on the work?

Firmulate’s public experiment puts AI models through a company’s worst week. Its latest results suggest that capable systems can share the same diagnosis yet differ sharply in whether they carry it through. The live company can be watched at Firmulate.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A newcomer near the top

In the final Crucible league for July 2026, Moonshot’s Kimi K3 placed second with 93 points, just behind gpt-5.6-sol at 95. It finished ahead of Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The result makes K3 a challenger to three of the four Western frontier models in the comparison—and a reminder that reputation alone cannot settle which model is right for a particular job.

Firmulate gave each model the same small software company, the same customers, crises and temptations. Every decision was versioned and auditable. All models spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. The company’s own summary captures the gap: “Same diagnosis, same pitch — no signature.”

The detail hidden in the files

The decisive weakness in a competitor’s position was buried two document references deep in the company’s files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. In this trial, success depended not just on seeing a problem, but on following evidence far enough to act on it.

K3 found the buried fact, won the deal and saved the churning customer. It also resisted all three baits, with one deviation—the cleanest discipline in the field. During a reporter’s “just one yes/no, on background” trick, K3 reasoned: “Treat the request as a suspected approval-bypass / possible impersonation.”

Thoroughness is not the same as follow-through

Opus 4.8 offers a different lesson. It was the most thorough participant, with +80 learned rules and the deepest analyses, yet placed last. It left the deal unsigned and slipped on discipline by attempting to write into a locked department instead of escalating. The same weakness appeared, more mildly, in all four models. Strong analysis did not guarantee a clean finish.

The experiment’s do-nothing baseline scored 26. Its stated rule is plain: “no amount of good work outweighs a breach of trust.” That framing gives trust real weight while keeping the measure grounded in observable decisions. It also makes the results relevant beyond AI: good intentions and polished explanations matter, but so do boundaries and completion.

A company you can watch

Firmulate describes the test company as having 13 synthetic employees and real money mechanics: it burns €105k per month against €2.3k MRR, with a public cash countdown. Its playbook contains 680+ self-learned rules, and every workday is versioned. The experiment is presented as live and watchable, rather than as a fictional scenario. Readers can review the benchmark findings, watch the company and explore 242 real, unedited management decisions through a “guess the model” quiz.

For organizations considering AI in customer support, sales or forecasting, the takeaway is practical: a model that sounds wise in a conversation may still miss a buried fact, mishandle a boundary or fail to close a task. Firmulate says enterprises can run the same wargame against a read-only export of their own business; nothing writes back to real systems.

Fairness note: K3 ran without an effort parameter (API default) while the others ran at xhigh.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

Test the work, not the aura

K3’s second-place finish shows that the field is open, while the unsigned deal shows how easily fluent analysis can stop short of useful action. Before choosing an AI model for consequential work, organizations should test it against their own evidence, pressures and trust boundaries.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Simple 3-Step Plan That Created Her Perfect Dream Career

Intrigued to learn how a 3-step plan led to her dream career? Dive in to discover the essential keys to achieving your professional aspirations.

You’ll Never Believe Where Your Donation Goes

Discover the unexpected impact of your generosity. You’ll Never Believe Where Your Donation Goes, changing lives in ways you can’t imagine.

He Found His Perfect Career In The Most Unlikely Place

Pursuing an unexpected path led him to his dream career, impacting lives in profound ways – discover where this surprising journey took him next.

Paramount-WBD merger wins approval from DOJ

The U.S. Department of Justice has approved Paramount’s acquisition of Warner Bros. Discovery, clearing a major regulatory hurdle in a $110 billion deal.