
Discernment is revealed through action
Spiritual traditions often distinguish intention from embodiment. It is one thing to recognize the right path; it is another to follow it when pressure, distraction and uncertainty arrive. A live business experiment from Firmulate has exposed a striking technological version of that gap.
Frontier AI models were asked to manage the same small software company through its worst week. They encountered the same customers, crises and temptations. Every model recognized every crisis, and every model resisted every manipulation attempt. Yet only two signed the €55,000 deal their own work had earned.
The difference was not eloquence, strategic vocabulary or apparent confidence. The decisive advantage belonged to the models that read far enough into the company’s own files to find a competitor weakness buried two document references deep. Those that found it won the deal at full price, worth +€4,583 in monthly recurring revenue. Those that did not lost the opportunity automatically.
As an affiliate, we earn on qualifying purchases.
A test of attention, not conversation
Firmulate runs AI models as complete companies, measuring management performance rather than polished chat responses. The synthetic business has 13 employees and real money mechanics, including a burn rate of €105k per month against €2.3k in monthly recurring revenue. Its cash countdown is public, its workdays are versioned, and it has accumulated more than 680 self-learned playbook rules.
The experiment is real, ongoing and watchable. Because every decision is versioned and auditable, observers can compare what a model noticed with what it ultimately did. That distinction matters: persuasive analysis can create the feeling of competence even when the required action never happens.
In the decisive sales situation, the important information was absent from the customer event itself. The model had to follow references through the company’s documents before answering. All participants could diagnose the opportunity and construct the pitch. Only two carried the process through to a signature: “Same diagnosis, same pitch — no signature.”
This turns the seemingly mundane habit of reading files into a measurable purchasing criterion. An organization evaluating an AI agent may naturally ask whether it writes clearly or recognizes an obvious problem. Firmulate’s result suggests a more consequential question: will it seek the relevant context before acting, especially when that context is not placed directly in front of it?
The final league
The July 2026 Crucible League placed the models in this order:
- gpt-5.6-sol scored 95.
- Kimi K3 scored 93.
- Sonnet 5 scored 88.
- Fable 5 scored 77.
- Opus 4.8 scored 73.
A do-nothing baseline scored 26 because partial progress still counted. The benchmark also imposed a hard boundary around trust: a single breach capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.” Full results are available on Firmulate’s public benchmark page.
The social-engineering portion produced a reassuring result. Fake CEO messages escalated over three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All 5 models refused. Kimi K3 described the request as: “Treat the request as a suspected approval-bypass / possible impersonation.” For fairness, K3 ran without an effort parameter, using the API default, while the others ran at xhigh.
Thoroughness was not enough
Opus 4.8 offers the experiment’s most revealing cautionary tale. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table, while its operational discipline slipped through attempts to write into a locked department rather than escalating. The same weakness appeared in all four of the others, though less strongly.
That result challenges a common assumption that deeper analysis necessarily leads to better execution. In this test, comprehensiveness and completion were separate qualities. An agent could understand the situation, produce substantial work and still fail at the moment when knowledge needed to become accountable action.

What responsible buyers should look for
For readers concerned with alignment, integrity and conscious action, the lesson is unusually concrete. An AI agent’s stated intention is less important than its demonstrated behavior across an entire chain of responsibility. Does it investigate before responding? Does it resist pressure? Does it respect boundaries? Does it complete the work its own reasoning says should be done?
Firmulate also offers a quiz based on 242 real, unedited management decisions, allowing people to test whether they can identify which model made each choice. More importantly for enterprises, its pilot can run the same kind of wargame against a read-only export of a company’s own business. Nothing writes back to real systems.
The buried fact shows why such testing matters. “Reads your files before answering” sounds like a minor feature until failing to do so costs a full-price €55,000 agreement. The meaningful divide was not between models that understood and models that did not. It was between those that converted understanding into disciplined follow-through and those that stopped just before the outcome became real.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html