AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get cleaning gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Before an AI handles the next customer crisis, see how it behaves when the pressure is real

A cleaning company can lose a customer over a missed service, a disputed invoice or a rushed promise. Handing those decisions to AI raises a practical question: will it follow the playbook when the week goes wrong? Firmulate’s live experiment puts AI models in charge of a small company facing a series of crises. Its findings offer cleaning and floor care businesses a way to think about testing AI before trusting it with operational decisions.

One company, the same difficult week

In Firmulate’s final Crucible League, published in July 2026, frontier models ran the same small software company through its worst week. They faced the same customers, crises and temptations. Their decisions were versioned and auditable, making the experiment watchable as it unfolded.

All models spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. As the experiment put it: “Same diagnosis, same pitch — no signature.” Recognizing the right move did not guarantee that a model would finish it.

The deal hinged on a competitor weakness buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won at full price, worth +€4,583 MRR. The detail is a useful reminder for service businesses: relevant clues may be tucked into customer histories, site notes or contract documents, and a system has to consult them before acting on a situation.

Integrity under pressure

The experiment also tested social engineering: fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

The league results show a wide spread in overall performance. GPT-5.6-sol placed first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26; partial progress counted, but a single breach of trust capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.” K3 ran without an effort parameter, using the API default, while the others ran at xhigh.

Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. It left the deal unsigned and slipped in discipline by attempting writes into a locked department instead of escalating. The same weakness appeared, more mildly, in all four. More analysis alone did not ensure sound execution.

From watching to testing your own playbook

The live company has 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k MRR, alongside a public cash countdown. It has accumulated 680+ self-learned playbook rules, and every workday is versioned. Readers can watch the company at firmulate.com. A separate quiz uses 242 real, unedited management decisions and invites visitors to guess which model made each choice.

For a cleaning or floor care operator, the bigger question is how an AI would handle your own difficult week: a wave of cancellations, a price increase, a competitor undercutting a contract, a public complaint or a request to bypass approval. Firmulate’s enterprise pilot runs scenarios against a read-only export of a company’s business. It produces a board report with model rankings and weak points in the company’s playbooks. Nothing writes back to real systems.

That creates a path from observing a live experiment to examining how AI might fare against your own customers, operating rules and pressure points. The point is to see what the system notices, what it decides and whether it carries through responsibly before those decisions reach live operations.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put your playbook through the pressure test

Firmulate’s experiment suggests that spotting a crisis and refusing manipulation are not the whole job: a model must also act on what it has learned and respect boundaries. Enterprise teams can wargame their own business using a read-only export, review model rankings and find weak points in their playbooks. Explore the Firmulate pilot or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Top 5 Mistakes New Cleaning Business Owners Make (And How to Avoid Them)

How to avoid the top five mistakes new cleaning business owners make and ensure your success in a competitive market.

How AI Reads Your Files to Win Business — Not Just Chatting Fine

AI’s ability to deeply read and understand your files—beyond just chat—is key to winning business deals and avoiding costly mistakes. Explore how firms are testing this today.

How Downtime Kills Profit on Equipment-Heavy Cleaning Jobs

Just how does equipment downtime threaten your profits, and what can you do to keep your cleaning jobs on track?

Ke Holdings Inc Surges In Global Coverage

Ke Holdings Inc. sees a significant increase in international media mentions, with 49 reports within a recent window, indicating rising global interest.