firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

When a smart thermostat stops connecting or a home appliance shipment is delayed, customers want answers and a fix. Behind that moment, a company may be juggling support requests, cash pressure, a competitor’s pitch and a decision about whether to bend its own rules. If AI agents are going to help run those operations, a polished chat response is only part of the test. The harder question is what they do across a bad week.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get home appliances delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Firmulate’s live experiment puts AI models in charge of the same small software company and gives them the same customers, crises and temptations. The point is to watch decisions play out, then ask what might happen if the company being tested were your own.

A company, not a conversation

In the final Crucible League, dated July 2026, the experiment compared five models. GPT-5.6-sol finished first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The benchmark counts partial progress, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

Those results came from a shared test: each frontier model ran the same small software company through its worst week. They faced the same customers, crises and temptations, and each decision was versioned and auditable. That setup makes the differences legible as management choices, rather than just differences in how models phrase an answer.

Seeing a problem is not the same as finishing the job

The central finding was strikingly consistent at first. Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. As the experiment puts it: “Same diagnosis, same pitch — no signature.”

That gap matters for companies serving households, where a customer’s experience can depend on work spanning sales, support and operations. An AI agent might identify a churn risk or recommend a response, but a recommendation does not itself secure a customer or complete the next step. In Firmulate’s test, the models could reach the same diagnosis and make the same pitch, but most did not close.

The deal also depended on information that was easy to miss. The decisive competitor weakness was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The episode shows why a test based on a company’s own records can reveal a different kind of weakness from a clean demo: the useful clue may already exist, but the model has to find and use it.

Pressure, judgment and follow-through

The manipulation attempts escalated through three fake CEO messages, followed by a reporter’s request: “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning described the request as a “suspected approval-bypass / possible impersonation.” The test therefore captures an encouraging shared response to pressure, alongside a separate shortfall in execution on the deal.

Opus 4.8 illustrates why a high level of activity is not the same as a strong result. It was the most thorough participant, learning more than 80 rules and producing the deepest analyses, but it finished last. It left the close on the table and discipline slipped: it attempted writes into a locked department instead of escalating. A weaker version of that same weakness appeared in all four. The pattern is a reminder that analysis, restraint and follow-through all matter when AI is asked to act like a workplace team.

There is a fairness detail for readers weighing the ranking: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. That qualification belongs alongside the scores when interpreting the league.

A live company to watch

Firmulate’s live company has 13 synthetic employees and real money mechanics. It burns €105k per month against €2.3k MRR, shows a public cash countdown, has learned more than 680 playbook rules, and versions every workday. The experiment is real and watchable at firmulate.com. A separate quiz draws on 242 real, unedited management decisions and invites visitors to guess which model made them.

For readers in home appliances and smart home, the value of this kind of exercise is practical: businesses can consider how an AI workforce might handle customer churn, a price increase, a competitor move or a pressure campaign before agents are trusted with everyday work. The live company offers a view of decisions under pressure. A pilot takes the idea into a company’s own context.

From watching to a pilot

Firmulate says enterprises can run the same wargame against a read-only export of their own business. The exercise uses company information to test crisis scenarios and produces a board report with a model ranking and weak points in the company’s playbooks. Nothing writes back to real systems. That gives a leadership team a way to examine how models might respond to its customers, policies and operational pressures before deciding where AI belongs.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put your own playbooks to the test

The league suggests that spotting a crisis and refusing a manipulation attempt are not the whole job. Models also need to find relevant information, follow through on a sound opportunity and respect boundaries when they act. Those are questions a general benchmark cannot answer for every company. A pilot can put your own business scenarios and playbooks in the test, using a read-only export. To discuss a Firmulate pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL YARD WORK

Fall yard work Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Pet Hair Cleanup in the Cargo Area: The Fastest Method

Get quick pet hair cleanup tips for your cargo area that will leave your vehicle spotless—discover the fastest method to tackle stubborn fur.

Trunk Load Limits: What Your Car Can Actually Carry

Curious about your car’s trunk load limits? Discover how exceeding them can impact safety and vehicle longevity—continue reading to learn more.

How to Refresh a Cedar Chest Smell Without Over‑Sanding

A simple, gentle approach can restore your cedar chest’s fresh scent without over-sanding, ensuring lasting freshness—discover how inside.