Case study · Unsolicited audit · B2B manufacturing · August 2026
A certified manufacturer with decades in its category. Not a client — a company in a category worth measuring, audited at full report depth and shared here anonymized. Every number below comes from a scored run, recorded verbatim with dates, and survived four verification rounds before publication.
Thirty prompts built across five intent categories — category discovery, brand-direct, comparison, problem-first, and trust. Three engines: ChatGPT, Gemini, and Claude. Clean sessions, no history, no personalization. Every answer recorded verbatim with its citations and its date. Competitors scored on the identical prompt set, one count per scored answer, with prompted mentions separated from earned ones.
That last distinction — prompted versus earned — is the entire audit. Hold onto it.
Appearance rate — named in 16 of 30 answers. ChatGPT 5 of 10, Gemini 5 of 10, Claude 6 of 10. Reads like a healthy number. It isn't one yet.
Share of answer — 16 mentions out of 130 vendor mentions across the same thirty answers. The engines are talking about this category constantly. Mostly about other people.
Citation share — of the 25 answers that showed citations, 11 cited a company-owned page. The other 14 built their answer from third parties.
The headline number. 16 of 30 answers named the company. Every engine contributed — which is exactly why it reads like a marketing win.
Named in half the answers on every engine, within six points of each other. Consistency across three engines usually means the cause is structural, not one model's quirk.
Sixteen appearances. Split them by how the engine got there, and the marketing win evaporates:
By intent, the same shape: brand-direct 100%, comparison 50%, trust 50%, category discovery and problem-first both 33.3%. The company owns the questions that contain its name and loses the questions that create customers.
Read it right to left. The company wins every question that already contains its name — and loses two out of three of the questions a stranger asks. The tall bar is the one that generates no new customers.
Meanwhile the category leader was named thirteen times on the identical prompt set — ten of them earned.
Strangers asking the category's questions almost never got this company's name back. People who already knew it always did.
Then a routine buyer question: is this company legitimate?
Two of three engines answered with an anti-fraud script — advising the buyer to use a credit card rather than a wire transfer for purchase protection — instead of the company's own credentials. One engine vouched without hedging. A coin-flip on the most basic question a procurement team asks.
The sharpest version of the finding: one engine reached the company's own website, cited it — and ran the fraud script anyway. The credentials were on the page. The engine wanted corroboration the page couldn't provide.
And the controlled result that makes it actionable: the same engine, same account, same day, name supplied both times, answered differently depending on which sources it happened to reach — vouching when it landed on distributor pages, hedging when it landed on aggregators. Not three engines disagreeing. One engine, split by source coverage.
Your own site is not corroboration.
One more, because it shows what verification is for. Early scoring flagged the engines as getting the company's founding date wrong — two answers gave two different dates decades apart.
Verified against the company's own site: both dates were accurate. The technology traces to one decade, the current entity to another, and the company's own pages state both without connecting them. The engines weren't wrong. They were reporting two true things with nothing machine-readable linking them — so every answer picked one at random.
The remediation isn't a correction. It's making the timeline machine-readable — schema and a labeled history — so the engines stop having to choose.
That finding survived four verification rounds, one of which reversed it. It's in this case study because it got reversed: an audit that can't change its mind under evidence isn't measuring anything.
The engines built their answers from distributors, directories, and review aggregators far more than from the company's own pages. Trust broke where those sources were thin. Presence there — not more site copy — was the ranked first fix.
Organization schema, a labeled timeline, credentials stated where extraction reaches them — so two true dates stop reading like a contradiction and certifications stop being invisible to a trust question.
Comparison prompts split 50%, and no vendor owned them. The engines synthesize comparison answers from whoever wrote one. Nobody had.
Same thirty prompts, same scoring. The change between two dated runs is the only number that means anything.
This company never asked for this audit and never paid for it. Publishing its name alongside its gaps would be trading its reputation for my marketing — the exact move this practice exists to be the alternative to. The subject received the findings. The method is public. The numbers are real.
When audits run under engagement, results are published only with written permission, named and dated.
The snapshot is the same instrument at a tenth of the scope — one prompt set, three engines, one page inside 48 hours. It answers whether you have this company's problem.
I run a limited number each month, so there may be a short wait.
Takes about a minute. The more precise the last two answers, the sharper the report.
I'll only use this to run and send your snapshot. No list, no sequence you didn't ask for, and you can tell me to delete it at any time (see the privacy policy). I run a limited number each month, so there may be a short wait.