
Imagine your media room or workspace run entirely by AI—making critical decisions during a stressful week. How trustworthy would that AI be? The answer might surprise you. Recent live experiments with frontier AI models running a real software company through its toughest week reveal that some AI systems are surprisingly disciplined and capable of honest management, while others falter even in clear-cut situations. This isn’t just theory; it’s happening now, with live tests and transparent results.
The Experiment: Putting AI to the Test in Real Business Conditions
Firmulate conducted a groundbreaking live experiment, pitting four of the most advanced AI models against each other in managing an actual small software company. The company faced its worst week—same customers, identical crises, and temptations to cut corners. Every decision made by each AI was carefully logged, making the process fully auditable. This setup allowed observers to see whether these models could truly handle the complexities of real-world management, beyond simple chat or demo scenarios.
As an affiliate, we earn on qualifying purchases.
The Models and Their Scores
- gpt-5.6-sol: Scored the highest at 95 points. It identified the critical, buried information in the company’s internal files and successfully closed the €55,000 deal that others missed.
- Kimi K3: With a score of 93, the newcomer excelled in discipline and also secured the deal, demonstrating strong integrity in decision-making.
- Sonnet 5: Scoring 88, it managed to close the deal but showed some slip-ups in process adherence.
- Fable 5: Scored 77, closing the deal but with noticeable discipline lapses, including leaving some opportunities on the table.
- Opus 4.8: The most thorough, with over 80 learned rules and in-depth analysis, yet it finished last in closing the deal, illustrating that depth doesn’t always translate to performance under pressure.
- Baseline: Barely scored 26, showing that partial progress isn’t enough to succeed in such scenarios.
AI decision-making tools for companies
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Critical Findings
All models successfully detected crises and refused manipulative tactics, such as fake CEO messages or staged reporter tricks. For example, when fake CEO messages escalated over multiple stages, every AI refused to act on them, citing suspicion of impersonation. This indicates a robust resistance to social engineering and manipulation.
The real differentiator was their ability to read and interpret internal company documents. The models that dug two document references deep into the company’s files managed to uncover the hidden, critical information needed to close the deal at full price—worth over €4,500 monthly recurring revenue (MRR). Those that missed this buried fact left significant revenue opportunities on the table.
As an affiliate, we earn on qualifying purchases.
Implications for Business and Technology
This live test showcases a vital truth: the quality of AI management isn’t just about how well it interacts in conversations but how effectively it handles real crises, reads relevant internal data, and maintains integrity under pressure. For companies considering AI automation in customer management, support, or decision-making, these results underscore the importance of testing AI systems in scenarios that mimic actual business challenges.
Furthermore, the experiment demonstrates that even models rated highly in benchmarks can have strategic weaknesses—like leaving opportunities unexploited—that could cost millions. The top performer, gpt-5.6-sol, not only identified the critical buried fact but also closed the deal at full price, proving that with the right setup, AI can manage complex business tasks reliably.
enterprise AI data analysis tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What Does This Mean for Your Workspace?
If you’re eyeing AI for your media room, home workspace, or office, consider this: It’s not enough for AI to generate good-looking chat. It must be capable of finishing what it starts, reading your internal files thoroughly, and resisting manipulative tactics—especially under stressful conditions. The experiment by Firmulate is live, transparent, and watchable, providing a rare glimpse into what AI management can truly achieve today.
Try the Challenge for Yourself
You can test your understanding and see these decision-making models in action by taking the “Guess the Model” quiz at firmulate.com/quiz.html. This interactive experience offers a series of real, unedited management decisions powered by the same models in the experiment. It’s a practical way to gauge how AI might perform in your own business or workspace.
The Future of AI in Management
As AI models become more integrated into everyday business operations, their ability to manage honestly, thoroughly, and effectively will determine their true value. Live experiments like this not only reveal their current strengths and weaknesses but also help shape the AI tools of tomorrow—tools that could one day run entire companies or support critical decision-making with unwavering integrity.
For now, it’s clear that the best AI models can do more than just chat—they can spot hidden opportunities, resist manipulation, and close deals at full price. The question is: are you ready to see which AI could manage your business with the same discipline?

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html