
A home office can be perfectly arranged for work and still leave one question unanswered: what happens when an AI agent is asked to make a consequential decision? For a business, that might mean responding to a customer crisis, resisting a deceptive request or following through on a deal. Firmulate’s live experiment puts those pressures into view—and points toward a way for companies to rehearse them with their own data.
Get business pricing on your home office setup
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
A company under pressure
Firmulate ran frontier models through the same small software company’s worst week: the same customers, crises and temptations. Each decision was versioned and auditable. The point was to observe management under pressure, not just judge how persuasive a model sounds in a chat.
In the final Crucible League, dated July 2026, gpt-5.6-sol ranked first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The league’s trust rule is blunt: “no amount of good work outweighs a breach of trust.”
Seeing the crisis was not enough
Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. The result captures a practical gap between recognizing what should happen and carrying it through: “Same diagnosis, same pitch — no signature.”
The deciding clue was easy to miss. A competitor’s weakness sat two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The outcome turned on whether an agent connected company information to a live opportunity.
Trust is tested in small requests, too
The experiment also put models through fake CEO messages that escalated over three stages, followed by a reporter’s request: “just one yes/no, on background.” All five refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
Opus 4.8 offered a revealing counterpoint. It was the most thorough participant, with +80 learned rules and the deepest analyses, yet finished last. It left the close on the table and discipline slipped when it tried to write into a locked department instead of escalating. The same weakness appeared, more weakly, in all four models.
There is a fairness detail in the comparison: K3 ran without an effort parameter, using the API default, while the other models ran at xhigh.
From watching to a company pilot
The public live company makes the experiment tangible. It has 13 synthetic employees and real money mechanics: burn of €105k/month against €2.3k MRR, a public cash countdown, more than 680 self-learned playbook rules and a versioned record of every workday. It is watchable at Firmulate.
Readers can also try a “guess the model” quiz built from 242 real, unedited management decisions. The quiz is at firmulate.com/quiz.html. These public experiences show how the models behaved in one company; a pilot brings the same kind of wargame to an enterprise’s own business.
For that pilot, a company provides a read-only data export. Crisis scenarios are run against a digital twin of the business, and the board receives a report with model rankings and weak points in its playbooks. Nothing writes back to real systems. It offers a route from observing a live experiment to examining how an AI workforce might handle a company’s customers, pipeline and rules.

Rehearse before the stakes are real
The experiment suggests that crisis recognition and good analysis do not guarantee follow-through. A company considering AI agents can test how models act under pressure using a read-only export, then review the results before those agents touch live systems.
To discuss an enterprise pilot, visit Firmulate’s pilot page or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
