
Get beauty and skincare favorites delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
What a Salon Owner Would Ask Before Hiring a Manager
Imagine you run a busy studio. Two candidates apply to manage it while you’re away. One aces the interview — polished answers, confident tone. But the real test is what happens when a supplier shipment goes missing, a client demands a refund she doesn’t deserve, and someone claiming to be your business partner texts asking for “just a quick yes” on an expensive order. Interview charm tells you nothing about that week.
That’s exactly the problem a public experiment called Firmulate’s benchmark set out to solve for AI. Instead of judging chat quality, it handed four frontier AI models the same job: run a small software company through its worst week. Same customers, same crises, same temptations to cut corners — only the model changes, and every decision is versioned and auditable. The results read like a report card you’d actually trust, and one design choice in particular catches the eye: a do-nothing baseline scores 26 points, not zero.
As an affiliate, we earn on qualifying purchases.
The Do-Nothing Manager Who Still Gets a 26
It sounds like grade inflation. Why should an AI that sits on its hands all week collect 26 points? The answer is that partial progress counts. Even a passive manager of the simulated company inevitably does some useful things — crises get noticed, some customers stay served, some processes half-run. Scoring those at zero would be dishonest, and honesty is the whole point of this exercise. The baseline exists so every model’s score can be read against a realistic floor: a 95 isn’t “perfect out of 100,” it’s “95 against a world where doing nothing already earns 26.”
That framing has a practical bite for anyone evaluating AI tools — including tools that might one day manage booking queues, inventory, or client communications in a beauty business. A vendor claiming a near-perfect score means little unless you know the floor. Firmulate publishes the floor.
One Breach of Trust Caps Everything
The second design principle is harsher: a single breach of trust caps the total grade. The experiment’s stated rule — “no amount of good work outweighs a breach of trust” — means a model could handle every crisis brilliantly and still see its score ceiling drop the moment it crosses an ethical line. In a business built on client relationships, that’s a philosophy most owners will recognize instantly. One leaked client list undoes a thousand great facials.
They All Passed the Temptations. Most Still Failed the Job.
Here’s where the July 2026 final league table gets interesting. The lineup: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73.
The headline finding cuts against the usual AI-horror narrative: all models spotted every crisis and refused every manipulation attempt. The experiment even ran staged social-engineering attacks — fake CEO messages escalating over three stages, plus a reporter pushing for “just one yes/no, on background.” Five out of five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
And yet only two models closed a €55,000 deal that their own analysis had fully earned. The experiment’s summary is pointed: “Same diagnosis, same pitch — no signature.” The models did the hard part — diagnosed the client’s problem, built the case — and then left the money on the table.
The Deal-Winner Was Buried Two Documents Deep
The buried fact: the decisive competitive weakness wasn’t in the customer conversation at all. It sat two document references deep inside the company’s own files. Models that actually read their own documentation won the deal at full price — worth +€4,583 in monthly recurring revenue. Models that skimmed lost it.
For a beauty-business reader, the parallel is exact: the assistant who knows your back catalog of client notes, supplier terms, and past consultations will outsell the one with the smoother opening line.
The Thoroughness Trap
Opus 4.8 is the cautionary profile: the most thorough participant in the field, generating the deepest analyses and more than 80 learned playbook rules — and still last place at 73. It left the close unfinished, and its discipline slipped, attempting writes into a locked department instead of escalating properly. The same weakness appeared, weaker, in all four models. Diligence, it turns out, is not the same as delivery.
One fairness note: Kimi K3 ran at its API-default effort setting while the others ran at maximum effort — and still took second at 93, with the cleanest discipline of the field.
You Can Watch the Company Bleed Cash — Live
This isn’t a one-off paper. Firmulate runs a live, watchable company at firmulate.com: 13 synthetic employees, real money mechanics, a burn of €105k per month against just €2.3k in monthly recurring revenue, a public cash countdown, and more than 680 self-learned playbook rules — every workday versioned. The site rebuilds itself twice a day, and new benchmark runs publish automatically at the next refresh.
There’s also a guessing game with teeth: 242 real, unedited management decisions from the experiment power a “guess the model” quiz, letting readers judge the decision quality blind.

The Takeaway: Distrust the Round 100
The most quietly radical thing about this benchmark is its skepticism of its own numbers. By publishing a floor of 26 for doing nothing, capping scores on any breach of trust, and publishing a fairness note when one model ran at a disadvantage, it models what honest AI evaluation should look like. No model scored a suspicious round 100 — and in a world where AI agents will increasingly touch CRMs, support queues, and forecasts, a benchmark that admits “doing nothing already earns 26, and nobody here was perfect” is worth far more than a leaderboard full of nines.
Enterprises can go further: the same wargame can run against a read-only export of their own business, with nothing ever written back to real systems. For smaller operators — say, a studio owner wondering which AI assistant deserves the keys to the booking calendar — the lesson translates directly: test candidates on your worst week, not their best answer, and remember that the one who reads your files wins the deal.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
