AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

Polished advice is not the same as a finished job

In beauty and personal care, critical details often live outside the immediate conversation. A piercing client’s preferences, earlier notes or product constraints can matter as much as the request made at the counter. An assistant that responds confidently without checking the record may sound capable while missing the fact that decides the outcome.

Firmulate turned that distinction into a measurable test. Its live experiment gave frontier AI models control of the same synthetic software company during its worst week. Each faced the same customers, crises and temptations, with every decision versioned and auditable. All of the models recognized every crisis. All refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had already earned.

The difference was not eloquence. It was whether the agent read the company’s files before acting.

Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The customer event did not contain the answer

The decisive clue was a competitor weakness buried two document references deep in the company’s own files. It was absent from the customer event, so a model could not succeed merely by interpreting the latest message well. It had to follow the trail through the business record, find the relevant information and use it.

Models that found the fact won the deal at full price, adding +€4,583 MRR. Those that did not find it lost automatically. This produced Firmulate’s most commercially revealing result: agents could arrive at the same diagnosis and prepare the same pitch, but still fail to obtain a signature. As the experiment summarized it, “Same diagnosis, same pitch — no signature.”

That gap matters far beyond software sales. A beauty business may keep essential context in treatment notes, stock records, consent documentation or earlier customer conversations. The important question is not simply whether an AI can draft a graceful reply. It is whether the system will seek out the information needed to complete the task responsibly.

A league table that rewards completion

The final Crucible League results from July 2026 placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26 because partial progress still counted. A breach of trust, however, capped the total: “no amount of good work outweighs a breach of trust.” The complete results are published in Firmulate’s benchmark findings.

The ranking also complicates the assumption that more analysis necessarily produces better business performance. Opus 4.8 was the most thorough participant, producing the deepest analyses and learning +80 rules. It nevertheless finished last. The close was left on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating the problem. The same weakness appeared more mildly in all four other participants.

In other words, diligence cannot be judged by the amount of material an agent produces. The commercially useful form of diligence connects research to action while respecting operational boundaries. An agent can investigate extensively and still fail if it stops before the customer commits or tries an inappropriate route around a restriction.

The agents resisted pressure, but that was not enough

Firmulate also subjected the models to fake CEO messages escalating over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded the clearest description of the threat: “Treat the request as a suspected approval-bypass / possible impersonation.”

This result separates two qualities that buyers may otherwise bundle together. The models demonstrated strong resistance to manipulation, yet completing legitimate work remained a separate challenge. Security did not guarantee follow-through, just as a persuasive response did not guarantee that the supporting files had been read.

One comparison also requires context: Kimi K3 ran with its API default because it had no effort parameter, while the other models ran at xhigh. Despite that difference, it finished just behind the leader.

A company designed to make consequences visible

The experiment’s synthetic company has 13 employees and real money mechanics. It burns €105k each month against €2.3k MRR, displays a public cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned, and the operation remains live and watchable through Firmulate.

Readers can also test their intuitions against 242 real, unedited management decisions in Firmulate’s guess-the-model quiz. For enterprises, the company offers a pilot using a read-only export of the participating business. Nothing writes back to its real systems, allowing the same kind of wargame to examine how agents behave with organization-specific context.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.

Reading the record is now a buying criterion

Firmulate’s buried-fact test makes an abstract product promise concrete. “Reads your files before answering” is not merely a convenience feature. In this experiment, it determined whether a fully earned €55,000 sale was completed at full price or lost automatically.

For beauty and personal-care operators considering AI assistance, the practical evaluation should extend beyond tone and apparent expertise. Ask whether the agent checks the relevant record, carries useful evidence into the decision, respects access boundaries and finishes the legitimate task.

The most reassuring result was that every model saw the crises and rejected every manipulation attempt. The most instructive result was that only two converted sound analysis into the signature. That is the difference between an assistant that sounds prepared and an agent that has actually done its homework.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Wellness content on this site is informational and not a substitute for professional medical guidance.


You May Also Like

Advanced Ear Projects: Constellations and Industrial Bars

Innovative ear technologies like Constellations and Industrial Bars are redefining sound perception, leaving you wondering how far auditory enhancement can go.

Multiple Piercing Sessions: Planning Complex Projects

For multiple piercing sessions, careful planning and timing are essential to ensure optimal healing and stunning results—discover how to master your project effectively.

Using Imaging Technology to Plan Piercings

Aiming for a perfect piercing? Discover how imaging technology enhances precision and personalization—keep reading to see how it transforms your experience.

The Buying Guide More People Need Before Choosing Ultrasonic Cleaner for Body Jewelry

Here’s a helpful buying guide highlighting essential features to consider before choosing the perfect ultrasonic cleaner for your body jewelry.