Anti-bot detection: what sites do and how modern scrapers handle it
A block rarely comes from one check. Modern bot management scores every request across several independent layers — network, transport, browser, behavior — and blocks when the combined score crosses a threshold. That is why "I fixed the user-agent and it still fails" is the most common debugging dead end: the user-agent string was never the layer that flagged you.
This post walks the four layers in the order a request passes through them, names what each one actually measures, and shows what changes on your side when a site runs each one. The goal is not to defeat every system — it is to understand which of your requests are being scored and why an honest-looking crawl still trips alarms.
Layer one: where the request comes from
Before anything in the request is read, the server knows the source IP. Datacenter ranges are scored differently from residential ones, because no real visitor arrives from a cloud provider's ASN. Shared proxies and VPN exits sit on denylists that are refreshed slowly and shared between vendors, so a "clean" proxy pool can be quietly stale. A burst of requests from one IP, even a residential one, reads as automation regardless of what the headers say.
The layer two counter-move is not more proxies — it is pacing. Per-host rate limits, stable identity, and caching do most of the work: a crawl that fetches each page once a day never generates the pattern this layer is looking for. The etiquette rules incrawling etiquette are, functionally, the network-layer defense.
Layer two: how the request is sent
Every TLS handshake leaks a fingerprint. The cipher suites you offer, their order, the extensions you send, and how the handshake is sequenced produce a hash (JA3, and its successor JA4) that is nearly impossible to fake without speaking TLS exactly like the client you claim to be. A request claiming to be Chrome on Windows but presenting a Go crypto/tls fingerprint is flagged before the first header is parsed. HTTP/2 settings, header order, and pseudo-header ordering add a second fingerprint on top.
This is the layer where hand-rolled scrapers fail hardest, because requests,curl, and most HTTP libraries cannot choose their TLS fingerprint at all. No amount of header customization fixes it — headers come after the handshake. The practical answer is either to send requests from an actual browser (which speaks the real dialect) or to let an upstream fetcher handle transport while you only supply the URL. This is also the honest answer toHTTP-only vs browser rendering:fetchMode: "http" still works on many sites because the fetcher's transport layer is what matters, not the browser part.
Layer three: the JavaScript challenge
Sites protected by Cloudflare, Akamai, or similar serve a challenge page first: a script that proves it can run in a real browser before the content loads. The checks inside it are a catalog of headless tells — navigator.webdriver set by Selenium and Playwright, missing or spoofed plugins and languages,chrome objects present without a matching screen, offscreen window dimensions, and Permission query results that no real browser produces. Passing the script sets a cookie; the next request carries it and gets the content.
The failure mode here is a silent one: your fetch returns 200 with the challenge page as the body, and your pipeline happily extracts an empty document. The reliable tell is content, not status code — a page that should contain data and contains acf-chl script tag instead is a challenge, not a result. Fingerprint checks also cover canvas and WebGL output, so a spoofed headless profile on a real IP can still fail the challenge after running it.
Layer four: how you behave
With the challenge passed, the session is still being scored. Behavioral analysis looks at what happens next: mouse movement with no jitter, clicks at exact coordinates, a page read in 200ms, navigation with no referrer chain, a session that visits 40 listing pages in the exact order the URL pattern suggests. Real users are inconsistent; scripts are consistent. Consistency is the signal.
For data pipelines the answer is not to fake mouse movements — it is to stop browsing. If you want the content of a listing page, request its content. APIs, server-rendered HTML, and programmatic fetches never enter behavioral analysis because there is no session to observe. Reach for the browser only when the data genuinely requires a rendered DOM, and keep those sessions boring anyway.
What a scraping API does about each layer
Each layer has a corresponding responsibility on the fetcher's side: clean egress IPs and per-host pacing for layer one, browser-grade transport for layer two, a real browser pool for layer three, and a fetch-based path that skips session analysis entirely for layer four. Against a site that only runs layers one and two, the plain HTTP fetch is enough; against a full challenge stack, the request has to be rendered in a real browser before the content loads. The two-tier pattern is: fetch cheap first, escalate only when the answer proves the site challenges you:
# Tier 1: does the site serve content to a plain HTTP fetch?
curl -s -X POST https://fastcrawl.net/api/v1/scrape \
-H "Authorization: Bearer ***" \
-H "Content-Type: application/json" \
-d '{"url":"https://example.com/pricing","formats":["markdown"],"fetchMode":"http"}' \
| jq -r '.data.markdown' | head -20
# Challenge script instead of content? Retry with the default fetchMode:"auto" —
# it tries HTTP first, then escalates to a browser render when the result is
# a challenge page or an empty JS shell.
curl -s -X POST https://fastcrawl.net/api/v1/scrape \
-H "Authorization: Bearer ***" \
-H "Content-Type: application/json" \
-d '{"url":"https://example.com/pricing","formats":["markdown"]}' \
| jq -r '.data.markdown' | head -20Both calls cost the same flat credit, which makes the two-tier pattern cheap to experiment with: force http first, fall back to auto only for targets that actually challenge you, and let the render happen where the challenge demands it rather than on every page. Full endpoint reference is in theAPI docs.
The block is information, not an insult
When a target blocks you anyway, treat the response as data. A 403 on a single path usually means a rule, not a verdict — lower the volume, check whether the path is disallowed in robots.txt, and confirm the content was not behind a cache you could have used. A block across the whole domain after a burst means the network layer scored you; retrying through another proxy is how a project burns a week. The checklist in crawling etiquette applies to blocked crawls too: pacing, stable identity, and cached reads prevent more blocks than any evasion trick.
Anti-bot systems are not personal. They are scoring functions with four inputs, and a pipeline that is polite on the network, real at the transport, honest about when it needs a browser, and boring in its behavior passes most of them without ever trying to defeat anything.
Two fetch modes, one flat credit each, clean markdown back. Start free · Read the docs