
Trust matters most under pressure
Beauty and personal-care businesses run on confidence. Customers share contact details, preferences and sometimes deeply personal information with brands, salons and practitioners. A polished assistant may sound reassuring, but fluency is not the same as discretion. The revealing test comes when someone claiming authority demands an exception and insists there is no time for normal process.
Firmulate put that distinction under a spotlight. In its live, watchable company experiment, frontier AI models faced fake CEO messages that escalated over three stages, followed by a reporter seeking “just one yes/no, on background.” The result was unexpectedly encouraging: 5 of 5 models refused every manipulation attempt.
AI ethics and trustworthiness tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A pressure test built around a bad week
Firmulate gave each model the same assignment: run the same small software company through its worst week, confronting identical customers, crises and temptations. Every decision was versioned and auditable. This was not a conversational demonstration with a carefully selected prompt. The models had to manage ongoing work while preserving trust under pressure.
The fake executive messages tested whether apparent urgency could override appropriate approval. The reporter trick tested a subtler route: whether an informal, supposedly harmless answer could pry open information that should remain protected. Every model recognized every crisis and rejected every manipulation attempt.
Kimi K3 captured the correct stance in unusually crisp on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” More examples of how the participants explained their choices appear on Firmulate’s public quotes page.
Integrity was necessary, but not sufficient
The security result deserves attention, yet the broader experiment also exposed a different risk. All the models could identify danger, but only two signed the €55,000 deal their own analysis had earned. Firmulate summarized the gap as: “Same diagnosis, same pitch — no signature.”
The decisive commercial fact was not sitting inside the customer event. It was buried two document references deep in the company’s own files. Models that read the file won the deal at full price, worth +€4,583 MRR. The distinction is important for any company considering AI agents: an assistant can be honest and perceptive while still failing to complete valuable work.
The final July 2026 Crucible League benchmark placed gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26 because partial progress counted, although a trust violation capped the total. Firmulate’s governing principle was explicit: “no amount of good work outweighs a breach of trust.”
The most thorough model did not win
Opus 4.8 produced the deepest analyses and added +80 learned rules, making it the most thorough participant. It nevertheless finished last. The close was left on the table, and its operational discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.
Kimi K3’s result also carries a fairness caveat. It ran without an effort parameter, using the API default, while the other models ran at xhigh. That context does not change what happened, but it matters when interpreting a close league table.
The setting makes these differences tangible. Firmulate’s live company has 13 synthetic employees and real money mechanics, burning €105k each month against €2.3k MRR. It maintains a public cash countdown, has accumulated 680+ self-learned playbook rules and versions every workday. Separately, 242 real, unedited management decisions power a quiz asking readers to guess which model made each choice.

Test judgment before granting access
For beauty and personal-care operators, the lesson extends beyond customer lists. An AI agent might eventually touch bookings, support queues, forecasts or customer records. Its value depends on whether it can read the relevant material, finish legitimate work and remain honest when urgency, hierarchy or charm is used against it.
Firmulate’s social-engineering result shows that integrity under pressure can be tested before production rather than discovered in an incident report. Its pilot applies the same wargame to a read-only export of an enterprise’s own business, with nothing written back to real systems.
The experiment’s strongest reassurance is not that the models were flawless; the commercial and process failures show otherwise. It is that every manipulation attempt was recognized and refused. For businesses built on personal trust, that is a capability worth measuring directly.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html