FIG. 00SpaSeen study
August 2026 · 20 metros · n=556What actually gets a med spa named by AI
We asked ChatGPT, Perplexity and Gemini the question a patient asks — “best med spa in {city}” — across 20 US metros, then measured 556 clinics from Google Places against who each engine actually named. Two variables predicted being named. Almost everything the industry recommends did not.
- Metros
- 20
- Clinics
- 556
- Engines
- 3
- Total spend
- $2.62
Perplexity and Gemini across all 20 cities; ChatGPT across 12, because it costs roughly fifteen times as much per call. Website reads are free, so every website variable was measured on all 556.
FIG. 01The answer
Two levers, and they belong to different engines.
There is no single thing that makes AI recommend a clinic. There are two, they do not overlap, and each one buys a different engine. That is a better result than one universal lever would have been, because it turns a vague number into two specific diagnoses with two different fixes.
| Lever | Buys you | Does nothing for |
|---|---|---|
| Google review count | ChatGPT | Perplexity, Gemini |
| How often your site names its city | Perplexity | ChatGPT |
FIG. 02Lever one — reviews
More Google reviews, more often named by ChatGPT.
Clinics ChatGPT named had a median of 266 Google reviews; clinics it did not name had a median of 175.5 (q=0.008). The same variable does nothing on the other two engines — q=0.97 on Perplexity, q=0.75 on Gemini. This is the most solid result in the programme: four replications on ChatGPT across three different question wordings.
What matters is the count. Not recency, not what the reviews say. Treatments named in review text was tested and is null on all three engines.
| Google reviews | Named by ChatGPT |
|---|---|
| 0–50 | 29% |
| 50–100 | 17% |
| 100–200 | 41% |
| 200–400 | 32% |
| 400–800 | 43% |
| 800+ | 58% |
Bumpy but climbing.
There is no number that gets you in. Any copy promising one is false — including ours. The honest sentence is that more helps, and it keeps helping.
FIG. 03Lever two — location language
Sites that say their city more often are named more often.
How many times the homepage says the name of its city. Named clinics said it 7 times to unnamed clinics’ 4 on Perplexity (q=0.0000), and the result survives every control we put on it — page length, and the correction described below.
| City mentions | Named |
|---|---|
| 0 | 14% |
| 1–2 | 17% |
| 3–5 | 31% |
| 6–10 | 35% |
| 11–20 | 45% |
| 21+ | 55% |
Every band beats the one below it. No sweet spot, no reversal — the cleanest dose-response in the programme.
On Gemini it is suggestive, not proven, and we will not call it replicated. Gemini showed the same direction (6 against 5, q=0.0008), but after the correction below it fell to q=0.067 on a quarter less sample. That may be power rather than absence. It is not a replication until it is re-run at full size.
The mechanism, stated as a guess rather than a finding: a patient asks “in Austin”. A page that never says “Austin” has nothing for half the question to match.
FIG. 04A leak we found in our own data
The first version of this study was wrong, and here is how.
Our name matcher required every distinctive word of a clinic’s name to appear in the answer. The question is “best med spa in Austin”, so every answer contains “Austin” — and a clinic called Austin Med Spa has nothing left but “austin” once generic words are stripped. It matched every answer automatically and scored as named 100% of the time. Sixteen clinics were pure false positives on exactly this.
The leak was correlated with the variable being tested, which is the worst kind: a clinic named after its city also says its city all over its website. It inflated the very finding it was being used to measure. We re-ran with place-named clinics dropped, and the sample fell from 556 to 369.
| After the correction | Result |
|---|---|
| City mentions, Perplexity | holds |
| Reviews, ChatGPT | holds |
| City mentions, Gemini | lost |
Perplexity q=0.0000 before and after; reviews q=0.008 → q=0.011; Gemini q=0.0008 → q=0.067, no longer significant.
We are publishing this because a study that never reports finding itself wrong is a study nobody checked.
FIG. 05What did not survive
Most of what sounds right is a page-length artefact.
The strongest-looking finding in the whole programme was treatment vocabulary — how many of 16 common treatments a homepage names. Named clinics covered 8, unnamed 5, q < 0.001 on both Perplexity and Gemini. It replicated across two engines. Then it was divided by page length and collapsed to q=0.97, with the unnamed clinics fractionally ahead. A longer homepage names more of everything by accident.
City mentions was put through the identical test and passed it. That is the entire difference between the two, and the only reason one is written up as a lever and the other as a lesson.
What else came back null?
- Review rating, and review recency. Only the count moves anything.
- Treatments mentioned in review text. Nothing on any engine.
- Google Business Profile categories. Right direction everywhere, significant nowhere.
- Credentials named on the site. One engine only — which is what a false positive looks like.
- Social platforms linked from the homepage. Page length again.
- Booking-platform profiles. Points the wrong way, and the market is too fragmented to build on.
- Site freshness. Untestable, not disproven: 41% of clinic sites claim an update within the week, because site builders regenerate the date automatically.
FIG. 06The lever we found and did not build
The biggest effect in the study is one we cannot honestly sell.
Clinics whose site names a doctor were named far more often — 49% against 21% on Perplexity, 26% against 8% on Gemini, p < 0.000001 on both. It is the largest clean effect in the programme.
It only wins questions that ask for a doctor. Every question in our monitored set uses a qualifier the same sweep found null — “best”, “top rated”, “most affordable”. So it cannot move the number we report, and recommending it would be selling a fix for a score that provably cannot respond to it.
The tempting version we are deliberately not writing: most US states require a med spa to have a medical director, and most sites never mention one, so the fix is a single page edit. That explains why it is cheap, not why it is worth anything.
FIG. 07What a clinic should take from this
Two things to count, and a lot of advice to ignore.
Ask your front desk for reviews, and make sure your homepage says the name of your city in the places a reader would expect it. Neither needs an agency. Then be sceptical of anyone selling the rest of the list — we tested most of it, and it did not hold.
None of this tells you whether AI is naming your clinic today. That is a measurement, not a rule of thumb, and it is what SpaSeen does: the same patient questions across ChatGPT, Claude, Perplexity and Gemini, on a schedule, with the findings that explain the result.
Method — 20 US metros, 556 med spas sampled from Google Places, question “best med spa in {city}”. Perplexity and Gemini on all 20 metros, ChatGPT on 12. Comparisons are Mann–Whitney U on medians; q-values are Benjamini–Hochberg corrected across the variables tested, and q is what is reported throughout, never a raw p. Figures quoted from SpaSeen’s own sweep, August 2026. Published 2026-08-26.