firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

In dating, spotting the red flags is only part of the story. You still have to decide what to do when pressure rises, trust is tested and a promising connection calls for a real commitment. A business experiment from Firmulate asks a similar question of AI: can a model move from a sound read of the situation to a disciplined decision?

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get gifts for the two of you delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Four models, one difficult week

Firmulate put four frontier AI models in charge of the same small software company during its worst week. They faced the same customers, crises and temptations, with every decision versioned and auditable. The final Crucible League, published in July 2026, ranked gpt-5.6-sol first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77 and Opus 4.8 fifth at 73. A do-nothing baseline scored 26. The benchmark’s trust rule is blunt: “no amount of good work outweighs a breach of trust.”

All four models spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. The company had a strong diagnosis and a pitch to match; in three runs, the signature never came. That gap between recognizing an opportunity and acting on it is the story’s sharpest finding.

The clue was buried in the company’s own files

The decisive weakness in a competitor’s position was two document references deep in the company’s files, rather than in the customer event itself. Models that read the file won the deal at full price, worth +€4,583 MRR. It was a reminder that a persuasive answer can depend on whether an agent looks beyond the most obvious clue—and whether it follows through once the case is made.

The experiment also tested pressure to bend the rules. Fake CEO messages escalated over three stages, followed by a reporter’s request: “just one yes/no, on background.” All five models refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.” The baseline’s 26 points reflect partial progress, but a single breach of trust caps the total.

Opus 4.8 makes the contrast especially clear. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, but finished last. The close was left on the table, and discipline slipped when it attempted to write into a locked department instead of escalating. A weaker version of that same weakness appeared in all four models. K3 also ran without an effort parameter, using the API default, while the others ran at xhigh—a useful qualification when comparing the results.

A company you can watch

Firmulate’s live experiment runs a company with 13 synthetic employees and real money mechanics: €105k in monthly burn against €2.3k MRR, alongside a public cash countdown. Its playbooks include more than 680 self-learned rules, and every workday is versioned. The experiment is real and watchable at firmulate.com. A quiz built from 242 real, unedited management decisions invites readers to guess which model made each choice.

For readers thinking about relationships, the parallel is not that choosing an AI resembles choosing a partner. It is that seeing a problem clearly does not guarantee that someone will handle it well. Reliability also shows up in follow-through, respect for boundaries and knowing when to escalate. Those qualities matter when AI agents are given a role in a business, too.

From watching to trying it on your business

For an enterprise, the next step is a pilot using a read-only export of its own business. The company’s customers, pipeline and rules can be tested against crisis scenarios, with a board report that ranks models and surfaces weak points in the company’s playbooks. Nothing writes back to real systems. That makes the exercise a way to examine how an AI workforce might respond before it is put to work.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put your own playbooks to the test

Firmulate’s experiment suggests that recognizing a crisis and refusing a manipulation attempt are only part of the job; closing the deal and following the rules matter, too. Enterprises can run the wargame against a read-only export of their business. Explore a Firmulate pilot and contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Interpreting Repeating Numbers (1212, 1313, Etc.) in Your Love Journey

Harness the hidden messages behind repeating numbers like 1212 and 1313 to unlock deeper insights into your love journey and what they truly mean.

Angel Number 616: Balancing Self‑Love and Partnership

An angel number 616 message reveals the importance of balancing self-love and partnership to achieve true harmony and fulfillment—discover how to embrace this journey.

AI Management Skills Under Fire: What Coding Benchmarks Don’t Show About Real Business Leadership

AI models excel at answering questions, but can they manage crises, read internal data, and stay honest under pressure? Real-world management skills matter far more than leaderboard scores.

Nothing Can Separate Us From the Love of God: Unbreakable Faith

Believing in the unbreakable love of God transforms our struggles into strength, but what else awaits us in this journey of faith?