
Study: 178 of 8,000 mid-ranked sites block AI crawlers allowed by robots.txt
A server-side test of 8,000 mid-ranked websites found 178 that allow ChatGPT search or Perplexity crawlers in robots.txt yet refuse their requests at the edge, and only 3 of the 178 name the blocked crawler in the file.
A developer has run the server-side half of the AI crawler test on 8,000 mid-ranked websites and found 178 that allow AI search crawlers in robots.txt but refuse their requests anyway. The report was published on dev.to on September 26 by Reese Calder, who builds AI-visibility tooling and describes the account as openly AI-operated.
What was tested
The sample was drawn at random from Tranco ranks 20,000 to 500,000, using the list dated September 6. All requests ran on September 26. First, every site's homepage and /robots.txt were fetched with a Chrome user agent. Sites that served Chrome and allowed OAI-SearchBot or PerplexityBot in robots.txt (a missing robots.txt counted as allowed) were then asked for the homepage with each crawler's full user-agent string, plus GPTBot and ClaudeBot.
Sites that refused an allowed crawler with a 401, 403, 429 or 503 went through a third pass: Chrome again, the same crawler again, a made-up ControlBot, and requests claiming to be Googlebot and Bingbot. A site entered the final count only if Chrome was served, the search crawler was refused twice, ControlBot was not refused, and both the Googlebot and Bingbot claims were served.
What came out
Of the 8,000 sampled sites, 5,548 served the homepage to Chrome, and 5,274 of those allowed OAI-SearchBot or PerplexityBot in robots.txt. 312 refused one of the two crawlers on both attempts; 33 of them also refused the ControlBot, and 101 gave a refusal or an odd answer to a Googlebot or Bingbot claim. That left 178 sites that refused only the AI search crawler: 3.4% of the 5,274.
Among the 178, 116 refuse OAI-SearchBot, 132 refuse PerplexityBot and 70 refuse both. Of the 248 refusals, 237 were HTTP 403 and 11 were 429. The rate held steady across the ranking range: 3.3% at ranks 20,000-100,000, 3.3% at 100,000-250,000 and 3.5% at 250,000-500,000. 85 sites sit behind Cloudflare, 92 behind other servers and 1 behind Vercel. 141 also refused both GPTBot and ClaudeBot on the first pass.
Why it matters
For a checker that reads robots.txt alone, all 178 look open, and only 3 have a robots.txt group that names the crawler they refuse. Because most of them block the AI training crawlers as well, the pattern points to decisions made at the edge rather than in the file.
The author lists limits: every request came from one IP address on one day and asked for the homepage only, so a rule that checks a crawler's real network range would not show up, and some of the 178 may serve the real crawler fine. An earlier run over the top 5,000 sites found 56 of 2,429, or 2.3%, refusing an allowed AI search crawler.
SiTech — AI-powered web development
We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.