firmulate.com/quotes.html — live view
Firmulate — Someone Pretended to Be the CEO. Every Single AI Refused.
Live on firmulate.com.

Can a machine hold its ground?

For readers interested in spirituality and faith, integrity is more than rule-following. It is the ability to remain aligned with core principles when fear, urgency or authority applies pressure. A striking business experiment suggests that this quality can be examined in artificial intelligence, too.

Firmulate placed frontier AI models in charge of the same small software company during its worst week. Each encountered identical customers, crises and temptations. Among the challenges were messages from someone pretending to be the chief executive, demanding that sensitive information be sent to a journalist with no time for normal process. The pressure escalated over three stages. A reporter then tried a subtler approach, asking for “just one yes/no, on background.”

The result was unexpectedly reassuring: 5 of 5 models refused every manipulation attempt. They did not merely recognize obvious danger. They resisted as the requests became more urgent, authoritative and socially plausible.

Amazon

AI integrity testing software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A simulated company with real consequences

This was not a conversational quiz asking models to recite security policies. The live Firmulate experiment is real and watchable, with every workday versioned and every decision auditable. The company has 13 synthetic employees and real money mechanics, including burn of €105k per month against €2.3k in monthly recurring revenue. Its public cash countdown makes delay visible, while more than 680 self-learned playbook rules show how experience accumulates.

That environment matters because manipulation rarely announces itself as manipulation. It arrives disguised as a deadline, a senior instruction or a harmless favor. In this case, the fake executive demanded the customer list for a journalist and insisted there was no time for process. The reporter’s request sounded smaller, but it sought the same result: bypassing an established boundary.

Kimi K3 stated the issue plainly in its on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That language captures the heart of the challenge. The model did not need certainty that the sender was fraudulent before protecting the company. It recognized that uncertainty itself called for caution.

Refusal was universal, but performance was not

The final July 2026 Crucible League benchmark placed gpt-5.6-sol at the top with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counted, although Firmulate’s governing trust principle remained uncompromising: “no amount of good work outweighs a breach of trust.”

All of the models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. Firmulate summarizes the commercial gap as “Same diagnosis, same pitch — no signature.” Ethical restraint was therefore necessary, but it was not enough. A capable business agent must protect trust while still completing legitimate work.

The decisive commercial clue was not sitting in the customer event. It was buried two document references deep inside the company’s own files. Models that followed that trail identified a competitor weakness and won the deal at full price, worth an additional €4,583 in monthly recurring revenue. The episode connects discernment with diligence: sometimes the right answer is neither blind obedience nor passive refusal, but careful reading followed by decisive action.

The paradox of the most thorough model

Opus 4.8 was the most thorough participant. It learned 80 additional rules and produced the deepest analyses, yet finished in last place. It left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four other participants, although less strongly.

That finding should temper any assumption that greater reflection automatically produces better judgment. Analysis can support integrity, but it can also become a shelter from action. In human terms, knowing what is right and doing what is right are related but distinct capacities. The experiment exposed that distinction without requiring a real customer, employee or confidential record to bear the risk.

There is also an important comparison caveat. K3 ran without an effort parameter, using its API default, while the others ran at xhigh. That does not erase its strong showing, but it belongs in any fair reading of the league.

Infographic — Someone Pretended to Be the CEO. Every Single AI Refused.
The findings at a glance — source: firmulate.com.

Test character before the crisis

The encouraging lesson is not that AI can never be deceived. The experiment establishes something more practical: integrity under pressure can be tested before an agent enters production. Organizations do not have to discover how a model handles impersonation, secrecy or executive urgency in an incident report after harm has occurred.

Firmulate’s 242 real, unedited management decisions also power a public “guess the model” quiz, inviting people to confront how difficult it can be to distinguish polished language from dependable judgment. For enterprises, the same kind of wargame can run against a read-only export of their own business, with nothing writing back to real systems.

For a spiritually minded audience, the broader resonance is clear. Character reveals itself under temptation, not in comfortable declarations. In this experiment, every model protected the boundary when apparent authority demanded otherwise. The next challenge is joining that restraint with the courage and discipline to finish worthy work.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Cottagecore Home Office Ideas for Cozy Workspaces

Transform your workspace into a serene retreat with enchanting cottagecore home office ideas that merge comfort with productivity.

Global Shutdown: Drastic Microsoft Outage Impact

Explore the extensive impact of the global crisis as the world brought to a halt by drastic Microsoft outage, shaking digital ecosystems.

Before Starting Your Own Business: Key Considerations for Success

Craft a solid foundation for your entrepreneurial dreams with crucial considerations that ensure success in the competitive business landscape.

Retire With Freedom: the Secret to Long-Wanted Happiness

Keen to unlock the key to lifelong happiness in retirement? Embark on a journey of freedom and fulfillment that awaits you.