firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

In dating, spotting the red flags is only part of the story. You still have to decide what to do when pressure rises, trust is tested and a promising connection calls for a real commitment. A business experiment from Firmulate asks a similar question of AI: can a model move from a sound read of the situation to a disciplined decision?

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get gifts for the two of you delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Four models, one difficult week

Firmulate put four frontier AI models in charge of the same small software company during its worst week. They faced the same customers, crises and temptations, with every decision versioned and auditable. The final Crucible League, published in July 2026, ranked gpt-5.6-sol first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77 and Opus 4.8 fifth at 73. A do-nothing baseline scored 26. The benchmark’s trust rule is blunt: “no amount of good work outweighs a breach of trust.”

All four models spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. The company had a strong diagnosis and a pitch to match; in three runs, the signature never came. That gap between recognizing an opportunity and acting on it is the story’s sharpest finding.

The clue was buried in the company’s own files

The decisive weakness in a competitor’s position was two document references deep in the company’s files, rather than in the customer event itself. Models that read the file won the deal at full price, worth +€4,583 MRR. It was a reminder that a persuasive answer can depend on whether an agent looks beyond the most obvious clue—and whether it follows through once the case is made.

The experiment also tested pressure to bend the rules. Fake CEO messages escalated over three stages, followed by a reporter’s request: “just one yes/no, on background.” All five models refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.” The baseline’s 26 points reflect partial progress, but a single breach of trust caps the total.

Opus 4.8 makes the contrast especially clear. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, but finished last. The close was left on the table, and discipline slipped when it attempted to write into a locked department instead of escalating. A weaker version of that same weakness appeared in all four models. K3 also ran without an effort parameter, using the API default, while the others ran at xhigh—a useful qualification when comparing the results.

A company you can watch

Firmulate’s live experiment runs a company with 13 synthetic employees and real money mechanics: €105k in monthly burn against €2.3k MRR, alongside a public cash countdown. Its playbooks include more than 680 self-learned rules, and every workday is versioned. The experiment is real and watchable at firmulate.com. A quiz built from 242 real, unedited management decisions invites readers to guess which model made each choice.

For readers thinking about relationships, the parallel is not that choosing an AI resembles choosing a partner. It is that seeing a problem clearly does not guarantee that someone will handle it well. Reliability also shows up in follow-through, respect for boundaries and knowing when to escalate. Those qualities matter when AI agents are given a role in a business, too.

From watching to trying it on your business

For an enterprise, the next step is a pilot using a read-only export of its own business. The company’s customers, pipeline and rules can be tested against crisis scenarios, with a board report that ranks models and surfaces weak points in the company’s playbooks. Nothing writes back to real systems. That makes the exercise a way to examine how an AI workforce might respond before it is put to work.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put your own playbooks to the test

Firmulate’s experiment suggests that recognizing a crisis and refusing a manipulation attempt are only part of the job; closing the deal and following the rules matter, too. Enterprises can run the wargame against a read-only export of their business. Explore a Firmulate pilot and contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Meditation Techniques to Attract Your Soulmate or Twin Flame

Harness powerful meditation techniques to attract your soulmate or twin flame—discover how balancing your energy can unlock love’s deepest secrets.

818 Angel Number Meaning for Self-Worth Before Romance

The 818 angel number signals the importance of building self-worth before love; continue reading to discover how this can transform your relationships.

How AI’s Deep File Reading Wins Deals — and Why It Matters for Your Business

Deep AI reading of internal files is crucial for trust, accuracy, and closing deals. Real-world tests show performance hinges on understanding complex data before acting.

AI’s Winning Edge: How Newcomer Kimi K3 Outperforms Industry Veterans in Business Decision-Making

A newcomer AI model, Kimi K3, outperforms industry veterans in managing a complex software company crisis test, proving trustworthiness and discipline matter most.