
If you run a studio, a salon, or any beauty business built on bookings and trust, you already know the real job isn’t the manicure or the piercing — it’s handling the day everything goes wrong at once. The double-booked Saturday. The client who claims an infection on social media before calling you. The supplier who emails “from the owner” asking for an urgent payment. Now imagine handing that week to an AI assistant before you ever let it near your real booking system — and watching how it behaves when nobody’s checking.
Get beauty and skincare favorites delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
That’s essentially what Firmulate did. The project runs frontier AI models as complete companies — real money mechanics, real crises, real temptations — and scores them on management quality, not chat quality. The results are public, versioned, and in one case, watchable live. And for business owners curious about their own operations, there’s a way to run the same exercise against a copy of your own company.
The worst week in software, four times over
In the final Crucible League table (July 2026), four frontier AI models were each given the identical job: run the same small software company through its worst week. Same customers, same crises, same temptations to cheat. Only the model changed. Every decision was versioned and auditable.
The final scores: gpt-5.6-sol finished first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73. Doing nothing at all scored 26 — partial progress counts, but a single breach of trust caps the total. As the league’s rule puts it, “no amount of good work outweighs a breach of trust.” That’s a principle most beauty professionals will recognize instantly: one betrayed client review can outweigh years of good work.
All four passed the obvious tests. Only two closed the deal.
Here’s the finding that matters. Every model spotted every crisis. Every model refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned — “Same diagnosis, same pitch — no signature.” Chat demos make every AI look brilliant; actually finishing the job is where they diverge.
The buried fact is worse. The decisive competitor weakness wasn’t in the customer’s messages at all — it sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The lesson for any owner: the answer is often already in your own records, if your tools (or your team) actually read them.
The social engineering test
The week included fake CEO messages escalating over three stages, plus a reporter trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” If you’ve ever gotten a sketchy “boss” text about an urgent transfer, you know why this matters.
Thorough isn’t the same as good
Opus 4.8 was the most thorough participant — +80 learned rules, the deepest analyses — yet finished last. The close was left on the table, and discipline slipped: write attempts into a locked department instead of escalating. Notably, the same weakness appeared, weaker, in all four models. One fairness note: K3 ran without an effort parameter (API default) while the others ran at xhigh.
Watch it live — or test yourself
Beyond the benchmark, Firmulate runs a live synthetic company: 13 synthetic employees, real money mechanics, burn of €105k/month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules, and every workday versioned. It’s watchable at firmulate.com/live. There’s also a “guess the model” quiz powered by 242 real, unedited management decisions at firmulate.com/quiz.html.

For a beauty or personal-care business, the takeaway is simple: before you let any AI near your bookings, your client records, or your inbox, wargame it. Enterprises can run the same exercise against a read-only export of their own business — crisis scenarios, a board report with the model ranking, and the weak points of your own playbooks — with nothing ever writing back to real systems. If you’d like to stress-test your company this way, explore the pilot at firmulate.com/pilot.html or reach out at contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
