FIG. 00SpaSeen survey
12 metros · n=213One in eleven med spa websites cannot be read at all
Before an AI assistant can decide whether to recommend a clinic, it has to be able to fetch the clinic’s website. We tried that on 213 med spa sites across 12 US metros. Nearly one in eleven would not let us read them — and most of those sites were not broken. They were working perfectly for people, and refusing machines.
- Sites tested
- 213
- Read fine
- 91.1%
- Refusing
- 7.0%
- Broken
- 1.9%
FIG. 01The finding
Most of the failures are a setting, not a fault.
Only 1.9% of the sites were genuinely broken — DNS failures, dead hosts, redirect loops. The larger group, 7.0%, answered our request and turned it away. Fifteen sites in total, and the refusals were concentrated in two places: eleven behind SiteGround, four behind Cloudflare.
That concentration is the useful part. These are not fifteen clinics who each made a decision about AI crawlers. They are fifteen clinics whose hosting or CDN shipped bot protection switched on, and the protection does not distinguish between a scraper and the reader that decides whether ChatGPT can quote them.
A clinic in this group is invisible to AI for a reason that has nothing to do with its marketing, its reviews or its website copy — and nothing in any report it receives would ever mention it.
FIG. 02What we actually measured
Our reader, refused. Not ChatGPT, refused.
This distinction matters and it is the one most easily lost. We measured our own automated reader being turned away. We did not observe OpenAI’s crawler being turned away. It is the same class of block arriving at the same door, and a site refusing well-behaved automated readers in general is very unlikely to be making an exception for the AI ones — but it is not the same claim, and we will not write it as one.
The four assistants run 9 crawlers between them, and the split that is easiest to get wrong is between the ones that build a search index or fetch a page live while answering, and the ones that only gather training data. Blocking a training crawler costs a clinic nothing today. Blocking one of the other 6 is what makes it uncitable:
| Crawler | What it does |
|---|---|
| OAI-SearchBot | search index |
| ChatGPT-User | live fetch |
| PerplexityBot | search index |
| Perplexity-User | live fetch |
| Claude-SearchBot | search index |
| Claude-User | live fetch |
The crawlers whose refusal costs visibility. The training crawlers — GPTBot, ClaudeBot, Google-Extended — are not on this list: blocking one has no bearing on whether an assistant cites the clinic today.
FIG. 03What is wrong with this survey
Some of that 7% is our own fault, and we can prove it.
SiteGround escalated from a soft challenge to a hard 403 when the survey ran eight requests in parallel. Which means the refusal rate is partly a function of how we behaved, not only of who hosts the clinic. A politer crawler, or the same crawler from a different IP, would likely measure a lower number.
We are publishing that because it is true and because it cuts both ways: a real AI crawler arrives from a known IP range with a published user agent, and may be allowed where we were not. The honest reading of this survey is that roughly one med spa site in eleven is configured in a way that refuses automated readers — not that exactly 8.9% of med spas are invisible to ChatGPT today.
FIG. 04How to check your own
How do I check whether my site blocks AI crawlers?
It takes about a minute, and you do not need us for it.
Open yoursite.com/robots.txt in a browser. If it does not load, or you are shown a security challenge, that is itself the finding — a crawler meets the same wall. If it loads, read it for a Disallow under any of the crawler names above.
A clean robots.txt is not the end of it: bot protection at the CDN blocks requests before robots.txt is ever consulted, and that is the failure the survey found most of. Diagnosing that means fetching the page as a crawler and comparing it with what a browser gets, which is one of the checks SpaSeen runs.
Method — 213 distinct med spa websites across 12 US metros, sampled from Google Places. Each homepage fetched once by SpaSeen’s reader; a site counts as refusing when the host answered and declined, and as broken on DNS failure, connect timeout or redirect loop. Concurrency up to 8 requests, which is itself a limitation — see FIG. 03. Published 2026-08-26.