Crawling etiquette: robots.txt, rate limits, and being a good citizen
Every scraping project moves through the same four steps. You pick a target, you write the fetch, it works, and then you scale it. The trouble is step four. A script that reads one page a day is invisible. The same script pointed at 20,000 URLs with fifty threads is an incident on someone else's side of the wire, and incidents get answered with blocks.
Etiquette is not sentimentality. It is the set of choices that determines whether your pipeline still works next month, after the site operator notices you. The rules are mostly mechanical: read the file that says what the site wants, pace yourself below the level that triggers a WAF, identify yourself so a block is a conversation instead of a silent 403.
What robots.txt actually obliges you to do
robots.txt is not a law. It is the file where a site states its preference, and the preference is what separates "we would rather you didn't crawl the search pages" from "leave the site alone". The parsing rules are simple enough to hold in your head: group rules by user-agent, apply the most specific group that matches your identity, and inside that group the longest matching path wins, with Allow beatingDisallow when two rules are the same length. No matching rule means allowed.
Two mistakes cost people access. The first is treating Crawl-delay as the speed limit: Google ignores it, most engines ignore it, and no one enforces it on your behalf. You have to choose a rate yourself. The second is reading the wrong group. Rules are per user-agent, so a User-agent: * section is your rules only if no more specific group matches the identity you send. A single Disallow: / under every named bot means the operator is telling crawlers to leave; overriding that is how a five-minute job turns into a rewrite.
Inspect before you build. Read the target's rules once for the user-agent you actually send: the robots.txt checker prints the file and tests any path against it, and the raw robots.txt is a scrape away if you would rather parse it yourself. That is a thirty-second check that prevents the most common category of wasted crawl.
Rate limits are yours to set, not the site's
Most sites never publish a rate limit, and the ones that do publish it for your benefit, not theirs. So the number that matters is the one you pick: requests per second per host, not globally across your crawl. Twenty hosts at two requests per second each is a polite workload. One host at forty per second is an attack, and it reads like one in the access log even when your intent is a product feed.
A working default: one request every 1.5 seconds against a small or independent site, and five to ten concurrent requests against a large CDN-backed one that already serves search engines at volume. Speed up only after an on-time run. The asymmetry is what makes this worth the discipline: a slow crawl costs you a few extra minutes, and a block costs you the domain, plus the next week you spend finding out which of your requests caused it.
Repeat visits deserve more care than first visits, because you are now a recurring pattern in their logs rather than a one-off. A daily monitor that re-fetches one page is fine almost anywhere. A nightly crawl of the whole catalogue is not, unless the site is yours or you asked. This is the same argument asmonitors over polling: watching five hundred pages for change should cost five hundred fetches a day, not five hundred fetches an hour.
Identify yourself, and keep the identity stable
Send a user-agent that names you and a URL that explains you, something likeYourBot/1.0 (+https://yourdomain.com/bot). Rotating random desktop browser strings across a crawl is the pattern that gets challenged, because a burst of identically branded but differently fingerprinted requests is the signature every bot-management vendor is built to catch. A named crawl with a contact URL is one an operator can rate limit, allowlist, or email you about.
Reputation is attached to your identity, and it works in both directions. A blocked datacenter IP does not unblock itself in an hour; denylists are shared and they age slowly. Keeping one stable identity per target is what lets a site operator escalate with you instead of silently refusing you.
The cheapest etiquette of all is not fetching at all. Fastcrawl caches scrapes for 48 hours by default, so repeat requests inside that window never touch the target, and themaxAge parameter is how you decide when freshness is worth a real request. That is the lever from planning a scraping budget, and it doubles as a politeness mechanism.
The three-step walk, with pacing built in
Read the rules, size the job, then walk it at a rate you chose. Step one needs no key at all, so you can check any target before signing up for anything:
# 1. What does the site want? (guest endpoint, no key)
curl -s -X POST https://fastcrawl.net/api/v1/try \
-H "Content-Type: application/json" \
-d '{"url":"https://example.com/robots.txt"}' | head -c 600
# 2. Size the job before you spend anything on it
curl -s -X POST https://fastcrawl.net/api/v1/map \
-H "Authorization: Bearer spd_..." \
-H "Content-Type: application/json" \
-d '{"url":"https://example.com","max_urls":500}' > urls.json
jq '.links | length' urls.json
# 3. Walk it paced: one host, one request at a time
jq -r '.links[]' urls.json | while read -r u; do
jq -n --arg u "$u" '{url:$u, formats:["markdown"]}' \
| curl -s -X POST https://fastcrawl.net/api/v1/scrape \
-H "Authorization: Bearer spd_..." \
-H "Content-Type: application/json" --data @- >> pages.ndjson
sleep 1.5
doneStep three is deliberately serial. If the list is long enough that this is too slow, raise the sleep rather than the concurrency first, and switch toPOST /api/v1/batch/scrape only when the host is large enough to absorb it. Batching changes your wall-clock time, not the number of requests the target sees, so it is not a politeness discount. A hundred-page walk is still a hundred requests.
When the list is link-driven and you would rather not manage the queue, crawldoes the walking for you with a bounded worker pool and URL canonicalization, and it returns a job id you poll. One caveat worth knowing before you point it at a small site: the default max_pages assumes a site that can take a hundred pages in a run. Lower it for anything small.
The etiquette checklist
Six rules cover almost every situation, and none of them cost you a feature:
- Fetch
robots.txtonce per target, for your own user-agent, and honour the longest matching rule. - Pick a per-host rate and write it down in the code, not in your head. Small sites get ~1 request per 1.5s.
- Send a real user-agent with a contact URL, and keep it identical across the crawl.
- Discover with
mapbefore you fetch, so you know whether the job is 50 pages or 50,000. - Default to cached reads (
maxAgeat its 48h default) and use monitors for change detection instead of re-crawling. - Treat a block as information. If a target blocks you, the next move is lowering volume or asking for access, not retrying through a proxy pool.
None of this is enforceable, which is exactly why it works. A pipeline that stays inside these lines keeps running quietly for years, and the operator who could have blocked you never has a reason to look. Full endpoint reference and the flat per-call credit rule are in the API docs.
Discovery, caching and monitors on one flat credit per call. Start free · Read the docs