Building a real-time data pipeline with a scraping API
"Real-time" is the most expensive word in a data pipeline, and most teams buy more of it than they need. The question is never can we get this page in a second. It ishow stale can this row be before the product breaks, and the answer is different for a price feed, a job board, and a competitor's changelog.
So a real-time scraping pipeline is not one architecture. It is a latency budget per data source, plus the cheapest fetch pattern that fits inside it. This is how to pick that pattern and how to prove you actually hit it, instead of assuming the number on the request is the number your users experience.
Start with the freshness budget, not the architecture
Write down, per source, how old a record is allowed to be. Three buckets cover almost everything. Seconds: a trigger has to fire on the same page view, e.g. you are mirroring a listing you republish. Minutes to hours: an order book, a volatile price, a stock indicator. Daily or weekly: a catalog, a job feed, a policy page that changes monthly.
The bucket decides the fetch pattern, because cost and load scale with cadence, not with cleverness. A daily change check on 5,000 URLs is one crawl. A five-minute check on the same set is 288 crawls a day against someone else's origin, which is how a working pipeline becomes a blocked IP. Once the budget is written down, the design is mostly subtraction: pick the rung that fits, and delete the machinery the faster rungs would have required.
Rung one: fan out in parallel instead of stampeding
The mistake at the low-latency end is 500 concurrent requests, one per URL, all fired at once. That gets you throttled by the target and, if you were hoping to reuse a rendered page, it also gets you no cache benefit: nothing is finished when the next request for the same page arrives. The parallel primitive you want is a batched fan-out with a bounded width, and the batch endpoint takes up to 50 URLs at a time, 10 wide:
# A 50-URL micro-batch: parallel 10-wide, one billable request,
# one cache write per URL, per-item error codes back.
curl -s -X POST https://fastcrawl.net/api/v1/batch/scrape \
-H "Authorization: Bearer $FASTCRAWL_KEY" \
-H "Content-Type: application/json" \
-d '{
"urls": [
"https://example.com/products/1001",
"https://example.com/products/1002",
"https://example.com/products/1003"
],
"formats": ["json"],
"maxAge": 0,
"fetchMode": "auto"
}' | jq '{total, succeeded, failed,
rows: [.results[] | {url, success, error_code}]}'Two details decide whether the fan-out is actually real-time. First,fetchMode: "auto" is right for mixed targets: it tries plain HTTP in about a tenth of a second and escalates to a browser render only when the page needs JavaScript, so a batch of static pages does not pay browser latency for pages that never needed it. Read when you actually need a browserbefore forcing it.
Second, maxAge: 0 forces a fresh fetch. The default cache window is 48 hours, which is exactly what you want for an hourly job and exactly what will make a "real-time" dashboard show yesterday's price. Be deliberate about it in both directions: a fresh fetch per URL costs a credit and hits the target, a cache hit is free and instant. On a five-minute cadence, a sane compromise is a short maxAge that most requests inside the window hit.
Rung two: stop polling, let the change tell you
Above the seconds bucket, polling is mostly waste. The expensive part of most pipelines is not the fetch, it is the executions where nothing changed and you still paid for bytes and parsing. Scheduled monitors invert that: they re-fetch on a schedule, hash the normalized content, and only tell you when the hash moves.
# Watch a page daily at 09:00 Asia/Hong_Kong and POST on change.
curl -s -X POST https://fastcrawl.net/api/v1/monitors \
-H "Authorization: Bearer $FASTCRAWL_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://example.com/pricing",
"schedule": "daily",
"timezone": "Asia/Hong_Kong",
"hour": 9,
"webhook_url": "https://api.example.com/hooks/fastcrawl"
}'
# -> {"success":true,"id":"..."}
# Fire it now while you build the receiver.
curl -s -X POST https://fastcrawl.net/api/v1/monitors/{id}/run \
-H "Authorization: Bearer $FASTCRAWL_KEY"The webhook body is deliberately small and typed: event ismonitor.run.completed, and data carriesmonitor_id, url, changed,content_hash, and the new markdown truncated at 50,000 characters. That is enough to route, deduplicate and diff without a second fetch, which matters because the webhook fires on change only. A quiet day sends nothing, so your receiver never has to distinguish "no changes" from "the job died".
The differentiator at this rung is what gets hashed. Naive pipelines hash the raw HTML, and then report a change every time a session token, a relative timestamp or a rotating ad slot appears on the page. The hash has to be computed over normalized content (extracted text, whitespace collapsed, volatile nodes dropped) or you get a pager that fires hourly and tells you nothing. If the timing details matter to you,monitors versus polling covers the cost math and the cadence edge cases.
Rung three: the sink is where pipelines actually break
Fetching is the easy half. The half that pages you at 3am is what happens after the bytes arrive, and three rules cover most of it.
Normalize at the boundary, not in the consumer. Convert to your storage schema in the same process that received the payload, so two services never disagree about what a "price" is. Make writes idempotent. Batch fan-out and webhooks both retry, which means at-least-once delivery, which means your insert must be an upsert keyed on something stable from the page (a product id, a posting id) rather than on arrival time.Store the hash. The content_hash from a monitor is the cheapest dedupe key you will ever get, and it lets you answer "did we already process this exact version" without re-diffing a text blob.
Treat failures as data too. Every error comes back classified, so branch on it instead of retrying blindly: target_dns_error, target_not_found andtarget_tls_error are terminal and will not heal in five minutes;target_rate_limited and target_server_error are worth a backoff retry; scrape_antibot_error means a challenge page was served instead of content and the right response is a different fetch mode, not another attempt. Each failure carries a retryable flag, and failed requests are not charged, so a retry loop that respects the flag costs nothing extra and stops hammering a dead host.
Measuring freshness, not claiming it
A pipeline is real-time when you can show the lag. Instrument two timestamps per record: when the page's own content says it was published or updated, and when your row was written. The difference is your true freshness, and it is almost always larger than the polling interval, because renders queue, retries back off, and the target's origin serves a cached copy it never tells you about.
Track p50 and p95 of that delta per source, not one number for the pipeline. A median lag of 40 seconds with a p95 of 20 minutes is a different product from a flat two minutes, and only one of them survives an incident. The same measurement tells you where to spend: if the lag sits in the fetch, lower maxAge or widen the batch; if it sits after the webhook fires, the bottleneck is your consumer and no scraping change will help.
One honest boundary. Scraping has no push channel: there is no socket a target site opens to tell you a page changed. Sub-second freshness on a source you do not control is a fiction, and any vendor promising it is selling you a request per second per page at someone else's expense. The achievable ladder is fan-out in seconds, monitors in minutes, and a measured lag you can quote.
Batch fan-out and scheduled monitors are both in the free tier. Start free · Read the docs