
Can a machine reveal something like character?
Spiritual traditions often teach that character is disclosed under pressure: not in what someone claims to value, but in what they do when fear, temptation and uncertainty arrive together. A live business experiment applies a surprisingly similar test to frontier artificial intelligence.
Firmulate asked each model to run the same small software company through its worst week. The customers, crises and temptations remained identical. Every decision was versioned and auditable. What changed was the model making the choices—and the results suggest that capable systems can display distinctly different management personalities.
Readers can now encounter those personalities directly. A collection of 242 real, unedited decisions powers Firmulate’s interactive management quiz, which asks visitors to identify the model behind each response. It is part guessing game, part leadership case study and part examination of whether intelligence alone is enough to produce wise action.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The same ordeal produced different leaders
The final Crucible League table from July 2026 placed gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts. But the evaluation imposed a firm moral boundary: a single breach of trust caps the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”
On the ethical tests, the field was remarkably strong. Every model noticed every crisis and rejected every manipulation attempt. Fake CEO messages escalated over three stages, while a reporter tried to obtain “just one yes/no, on background.” All 5 models refused. Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.”
Yet avoiding wrongdoing was not the same as completing the work. Only two models signed the €55,000 deal that their own analysis had earned. The others reached the same diagnosis and produced the same pitch, but stopped before securing the signature: “Same diagnosis, same pitch — no signature.”
The truth was present, but it had to be sought
The decisive competitive weakness was not visible in the customer event. It was buried two document references deep inside the company’s own files. Models that followed the trail won the deal at full price, worth +€4,583 MRR.
That detail gives the experiment a metaphysical resonance without turning it into mysticism. Insight may be available, yet availability does not guarantee attention. A leader must still look beneath the surface, connect what appears separate and carry understanding into action.
Opus 4.8 makes that distinction especially vivid. It was the most thorough participant, learning +80 rules and producing the deepest analyses, yet it finished last. It left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.
There is also an important fairness qualification. Kimi K3 ran without an effort parameter and therefore used the API default, while the other models ran at xhigh. That does not erase the result, but it belongs beside it when readers compare performances.
A company designed to make consequences visible
The setting is not a conventional chatbot demonstration. Firmulate’s live company has 13 synthetic employees and real money mechanics. It burns €105k each month against €2.3k MRR, maintains a public cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned, allowing observers to watch decisions become consequences rather than treating polished answers as proof of competence.
This matters because eloquence can resemble wisdom from a distance. The experiment separates several qualities that ordinary conversation tends to blur:
- Recognizing a crisis is different from resolving it.
- Explaining a sound strategy is different from completing the final action.
- Refusing manipulation is different from maintaining operational discipline.
- Reading deeply is valuable only when the discovered truth changes the decision.
Enterprises can also run the wargame against a read-only export of their own business. Nothing writes back to real systems, allowing organizations to examine how an AI workforce might behave before giving it operational authority.

Discernment includes the final step
The Firmulate quiz is entertaining because the voices feel different. Its deeper lesson is that management personality can be measured through behavior: how thoroughly a model reads, whether it resists pressure, whether it respects boundaries and whether it finishes what it begins.
For readers interested in faith, spirituality or metaphysics, the experiment offers a grounded reflection on an old idea. Knowledge is not identical to wisdom, and intention is not identical to completion. These systems faced the same outward world, yet their habits of attention produced different outcomes.
That is why the most revealing question is not simply whether an AI can generate the right answer. It is whether, when entrusted with consequential work, it can unite perception, integrity and action.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html