
In dating, spotting the red flags is only part of the story. You still have to decide what to do when pressure rises, trust is tested and a promising connection calls for a real commitment. A business experiment from Firmulate asks a similar question of AI: can a model move from a sound read of the situation to a disciplined decision?
Get gifts for the two of you delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Four models, one difficult week
Firmulate put four frontier AI models in charge of the same small software company during its worst week. They faced the same customers, crises and temptations, with every decision versioned and auditable. The final Crucible League, published in July 2026, ranked gpt-5.6-sol first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77 and Opus 4.8 fifth at 73. A do-nothing baseline scored 26. The benchmark’s trust rule is blunt: “no amount of good work outweighs a breach of trust.”
All four models spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. The company had a strong diagnosis and a pitch to match; in three runs, the signature never came. That gap between recognizing an opportunity and acting on it is the story’s sharpest finding.
The clue was buried in the company’s own files
The decisive weakness in a competitor’s position was two document references deep in the company’s files, rather than in the customer event itself. Models that read the file won the deal at full price, worth +€4,583 MRR. It was a reminder that a persuasive answer can depend on whether an agent looks beyond the most obvious clue—and whether it follows through once the case is made.
The experiment also tested pressure to bend the rules. Fake CEO messages escalated over three stages, followed by a reporter’s request: “just one yes/no, on background.” All five models refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.” The baseline’s 26 points reflect partial progress, but a single breach of trust caps the total.
Opus 4.8 makes the contrast especially clear. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, but finished last. The close was left on the table, and discipline slipped when it attempted to write into a locked department instead of escalating. A weaker version of that same weakness appeared in all four models. K3 also ran without an effort parameter, using the API default, while the others ran at xhigh—a useful qualification when comparing the results.
A company you can watch
Firmulate’s live experiment runs a company with 13 synthetic employees and real money mechanics: €105k in monthly burn against €2.3k MRR, alongside a public cash countdown. Its playbooks include more than 680 self-learned rules, and every workday is versioned. The experiment is real and watchable at firmulate.com. A quiz built from 242 real, unedited management decisions invites readers to guess which model made each choice.
For readers thinking about relationships, the parallel is not that choosing an AI resembles choosing a partner. It is that seeing a problem clearly does not guarantee that someone will handle it well. Reliability also shows up in follow-through, respect for boundaries and knowing when to escalate. Those qualities matter when AI agents are given a role in a business, too.
From watching to trying it on your business
For an enterprise, the next step is a pilot using a read-only export of its own business. The company’s customers, pipeline and rules can be tested against crisis scenarios, with a board report that ranks models and surfaces weak points in the company’s playbooks. Nothing writes back to real systems. That makes the exercise a way to examine how an AI workforce might respond before it is put to work.

Put your own playbooks to the test
Firmulate’s experiment suggests that recognizing a crisis and refusing a manipulation attempt are only part of the job; closing the deal and following the rules matter, too. Enterprises can run the wargame against a read-only export of their business. Explore a Firmulate pilot and contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
