firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

A home office can be perfectly arranged for work and still leave one question unanswered: what happens when an AI agent is asked to make a consequential decision? For a business, that might mean responding to a customer crisis, resisting a deceptive request or following through on a deal. Firmulate’s live experiment puts those pressures into view—and points toward a way for companies to rehearse them with their own data.

Buying for a business?Offer from Amazon

Get business pricing on your home office setup

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

A company under pressure

Firmulate ran frontier models through the same small software company’s worst week: the same customers, crises and temptations. Each decision was versioned and auditable. The point was to observe management under pressure, not just judge how persuasive a model sounds in a chat.

In the final Crucible League, dated July 2026, gpt-5.6-sol ranked first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The league’s trust rule is blunt: “no amount of good work outweighs a breach of trust.”

Seeing the crisis was not enough

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. The result captures a practical gap between recognizing what should happen and carrying it through: “Same diagnosis, same pitch — no signature.”

The deciding clue was easy to miss. A competitor’s weakness sat two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The outcome turned on whether an agent connected company information to a live opportunity.

Trust is tested in small requests, too

The experiment also put models through fake CEO messages that escalated over three stages, followed by a reporter’s request: “just one yes/no, on background.” All five refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Opus 4.8 offered a revealing counterpoint. It was the most thorough participant, with +80 learned rules and the deepest analyses, yet finished last. It left the close on the table and discipline slipped when it tried to write into a locked department instead of escalating. The same weakness appeared, more weakly, in all four models.

There is a fairness detail in the comparison: K3 ran without an effort parameter, using the API default, while the other models ran at xhigh.

From watching to a company pilot

The public live company makes the experiment tangible. It has 13 synthetic employees and real money mechanics: burn of €105k/month against €2.3k MRR, a public cash countdown, more than 680 self-learned playbook rules and a versioned record of every workday. It is watchable at Firmulate.

Readers can also try a “guess the model” quiz built from 242 real, unedited management decisions. The quiz is at firmulate.com/quiz.html. These public experiences show how the models behaved in one company; a pilot brings the same kind of wargame to an enterprise’s own business.

For that pilot, a company provides a read-only data export. Crisis scenarios are run against a digital twin of the business, and the board receives a report with model rankings and weak points in its playbooks. Nothing writes back to real systems. It offers a route from observing a live experiment to examining how an AI workforce might handle a company’s customers, pipeline and rules.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Rehearse before the stakes are real

The experiment suggests that crisis recognition and good analysis do not guarantee follow-through. A company considering AI agents can test how models act under pressure using a read-only export, then review the results before those agents touch live systems.

To discuss an enterprise pilot, visit Firmulate’s pilot page or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How I use HTMX with Go

A detailed overview of how developers are integrating HTMX with Go for dynamic web applications, including confirmed methods and best practices.

10 Home Deals From T.J. Maxx You Don’t Want To Miss This Labor Day

Discover the top 10 home deals from T.J. Maxx this Labor Day, offering significant savings on furniture, decor, and essentials. Limited-time offers not to miss.

How to Shop for Premium Sound in a Projector Room Without Overspending

Choosing affordable wireless speakers and DIY acoustic treatments can elevate your projector room’s sound—discover how to do it without overspending.

Adjusting Lip Sync Delays Manually

Optimize audio-visual sync by manually adjusting lip sync delays—discover the simple steps to achieve perfect harmony while ensuring your media plays flawlessly.