Back
A year of fighting scrapers on a 1.5-million-page website
SiTech Team3 წთ. საკითხავი

A year of fighting scrapers on a 1.5-million-page website

The operator of PatronView spent a year fending off automated scraping. In one week the server answered 2.5 million requests, while analytics recorded fewer than 6,000 human pageviews.

The philanthropy donor database PatronView, which holds about 1.5 million individual profile pages, has published a detailed account of a year spent fighting automated scraping. According to the site's operator, in the week the post appeared the server answered 2.5 million external requests and served 1.28 million full pages, while its analytics recorded only 5,977 human pageviews — roughly 214 bot page loads for every real one, and less than half a percent of total traffic.

The flood from China

The heaviest single event came on April 22, when the site took 3.6 million requests in one day from 361,844 unique IP addresses, almost all of them in China. Cloudflare's Managed Challenge absorbed 1.18 million of those in about ten hours, but a significant share passed through. On April 23 the operator blocked the whole country at the edge, later adding Vietnam and Singapore after similar waves. The site's genuine search traffic is 95.9% United States and 1.3% Canada.

The main metric: crawls per visitor referred

The post introduces the measure its author now applies to every decision: how many pages a crawler reads for each visitor it sends. Google needs 46 crawls per visitor, which he considers fair. In June, Anthropic's Claude-SearchBot requested 420,680 pages in a single week and delivered 12 human visitors — a ratio of 35,000 to 1, against the roughly 3,000 to 1 Cloudflare reports for Anthropic's crawlers. The bot consumed 4.63 GB of bandwidth; the humans it sent used 175 KB. Blocking it cut the crawler from 60,000 requests a day to about 25. Amazon's Amzn-SearchBot had meanwhile become the site's largest crawler at roughly 117,000 requests a day, and was blocked two days before publication.

Datacenters, home connections and results

In July the traffic changed shape: a steady wave of headless Chrome browsers on AWS, then Azure, while residential proxy networks routed requests through ordinary home connections and presented themselves as browsers from 2023. The countermeasures are deliberately blunt — challenge every continent except North America, challenge 46 datacenter ASNs, and challenge outdated browser versions. In a recent 48-hour window Cloudflare issued 106,437 challenges and only 252 were solved, a rate of 0.24%. In the first 24 hours after the final rules went live, 46,729 requests were blocked outright, 43,150 of them from Amazon's crawler. The site runs on Cloudflare Workers with D1 and KV; the normal bill is about $90 a month and bots account for 99% of usage. The author's conclusion is that scraping has become an economic problem — the fix he expects to work is pay-per-crawl — and his working rule is that a crawler which never sends a visitor gets blocked.

SSiTech

SiTech — AI-powered web development

We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.