Web scraping best practices for AI training data
A crawl is not a dataset. The gap between the two is every problem that shows up later: a model trained on scraped pages inherits the navigation menu as vocabulary, learns your boilerplate ten thousand times, memorizes a press release that appeared on forty syndicated domains, and quietly contains a page you were never allowed to keep. None of those failures announce themselves at ingest. They show up as flat eval curves and as a takedown email you cannot answer.
The fix is boring and it is all upstream: record provenance at fetch time, extract the document instead of the page, deduplicate before you filter, filter with rules you can audit, and refresh on a cadence you have priced. Every call below is a real request against fastcrawl.net, and the credit math is stated as you go.
Provenance at fetch time, not reconstruction
For every document you keep, store the URL as fetched, the canonical URL, the fetch timestamp, the HTTP status, the byte count, the fetch mode, and a content hash of the normalized text. That row is the answer to every question you will be asked later: which sources are in this checkpoint, when was this text last verified, why was this page excluded. Reconstructing any of it after the fact is impossible, because URLs rot, pages get edited, and the raw response is gone the moment you discard it.
Keep two copies: raw HTML in cold storage, normalized markdown in the hot set. A rule change (a new boilerplate filter, a different language threshold, a paragraph-level dedup you did not used to run) is then a reprocessing job over text you already have, not a re-crawl of ten thousand pages that may have moved. The content hash on the normalized text is also what makes cross-crawl dedup a join instead of a fuzzy matching project.
Decide the access question before you fetch, not after you train. Robots.txt is the floor, not the ceiling: honor disallow rules, skip anything behind a login or a paywall, and record the license signal you found (a rel="license" link, a Creative Commons footer, plain terms of service) next to the hash. Thecrawl etiquette rules cover pacing and identity; for a corpus, the part that matters is that the decision is stored per document, so an exclusion can be explained a year later.
Extract the document, not the page chrome
Raw HTML fed straight into a tokenizer wastes budget on chrome: headers, footers, cookie banners, related-article rails, "subscribe to continue" modals. Worse than wasted tokens, chrome is repeated n-grams. A template that repeats across 100,000 pages from 400 sites teaches the model the template, and repeated boilerplate is one of the documented ways a corpus degrades generation quality. Thewhat LLM-ready markdown means post has the four checks; the short version is that boilerplate stripping is the single highest leverage step in the pipeline.
Batch the fetch: up to 50 URLs per call, ten wide in parallel, one credit per URL, and a failed row reports error_code and retryable instead of failing the batch.
curl -s -X POST https://fastcrawl.net/api/v1/batch/scrape \
-H "Authorization: Bearer ***" \
-H "Content-Type: application/json" \
-d '{
"urls": [
"https://docs.example.com/guide/getting-started",
"https://docs.example.com/guide/install",
"https://blog.example.com/posts/launch-notes"
],
"formats": ["markdown"],
"fetchMode": "http"
}'
# -> {"success":true,"total":3,"succeeded":3,"failed":0,
# "results":[{"url":"...","markdown":"...","success":true}, ...]}fetchMode: "http" skips the browser entirely, which matters at corpus scale: static and server-rendered pages come back in about 100ms instead of the seconds a browser render costs, and theauto/http rule is what keeps a 50,000-page sweep inside a sane latency and credit budget. Store the markdown as the training text and the raw HTML as the reprocessing source. Both copies, one fetch.
Deduplicate before you filter
Exact duplicates are trivial and common: the same page reachable at three URL variants, the same article on the canonical domain and its www mirror, syndicated content appearing verbatim across outlets. Hash the normalized text and drop collisions. Keep the earliest fetch, and record which URLs lost so a source count stays honest.
Near-duplicates are where the real damage is. Wire services rewrite one clause, docs sites mirror a versioned subtree, and a forum thread is quoted back into itself. Shingle the text (5-word windows), compare with MinHash or an LSH band, and collapse anything above roughly 0.8 similarity to a single representative. Run this at paragraph level too: a boilerplate block that survived extraction and appears in most documents will otherwise be memorized as pattern, not content.
Dedup also runs across refreshes, and this is where caching earns its keep. A repeat scrape of a URL already fetched inside the 48-hour window is a cache hit and costs nothing, so a daily re-verification sweep of a corpus you crawled yesterday is free until the window expires. Thecredit and capacity math has the backpressure details; for dedup the practical rule is simpler: hash on ingest, join on the hash, and let the cache cover the short-interval re-checks.
Quality filters are rules, not vibes
Deterministic filters belong in code, before any model touches the text: minimum document length, language identification against your target languages, alphanumeric ratio, duplicate-line ratio, unmatched code fences, and truncation markers such as "subscribe to continue" or a partial-render notice. Each filter writes a reason code onto the row. Exclusion rate is the metric you will want to defend: if 18% of a crawl drops out, you need to say why, and "the filter said so" is not a reason.
For the judgment calls, use a schema-driven extract on the document you are keeping rather than an LLM reading your whole corpus. One call per candidate document pulls structured metadata you can filter on later.
curl -s -X POST https://fastcrawl.net/api/v1/extract \
-H "Authorization: Bearer ***" \
-H "Content-Type: application/json" \
-d '{
"url": "https://docs.example.com/guide/getting-started",
"prompt": "Describe this document for corpus filtering. Use null when the page does not make the answer clear; do not guess.",
"schema": {
"type": "object",
"properties": {
"language": { "type": ["string", "null"] },
"topic": { "type": ["string", "null"] },
"doc_type": { "type": ["string", "null"], "description": "article | docs | reference | forum | listing" },
"is_paywalled": { "type": "boolean" },
"has_pii": { "type": "boolean" }
},
"required": ["language", "doc_type", "is_paywalled"]
}
}'
# -> {"success":true,"data":{"language":"en","doc_type":"docs",
# "is_paywalled":false,...}} 1 credit per documentNullable fields stay nullable, for the same reason they do in price extraction: a missing answer must become null, a signal you can filter on, not a plausible value the model invented to satisfy the schema. One extract is one credit, failed calls are never charged, and this runs on the documents you are keeping, so it is a fraction of the fetch cost rather than a multiplier on it.
Refresh cadence and what it costs
Sources drift. Documentation gets rewritten, licenses change, a domain changes hands and the content with it. Re-fetch on a schedule you have priced rather than on demand: a full refresh of the corpus every quarter for static material, weekly for the volatile slice, and a hash monitor on the handful of sources whose terms or content actually move, so a change there triggers a targeted re-fetch instead of a full sweep.
The numbers are easy to plan. One credit per page fetched, free cache hits inside 48 hours, failed calls never charged. A 50,000-page corpus is 50,000 credits: the free tier's 1,500 credits a month cover a 1,500-document pilot that proves the pipeline end to end, and the $5 Go plan includes 5,000 credits with metered overage at $1.00 per 1,000 after that, so 50,000 fresh pages runs about $50 with nothing blocked mid-crawl. Compare that to a token-metered extractor billed per 1,000 characters, where the same corpus has a bill you cannot predict in advance, and seewhere the LLM-side savings actually come from.
The whole pipeline in one pass: gate on robots and license, fetch as markdown in batches, hash and shingle for dedup, apply deterministic filters with reason codes, enrich the survivors with one schema extract, and refresh on a priced schedule. Each stage is individually dull, and every stage is auditable after the fact. That is the property a training corpus actually needs: not that it is big, but that you can explain it.
Batch fetches, schema extracts and cache hits are one credit each, and failed calls are free. Start free · Read the docs