Notes on web data.
How AI agents read, extract and monitor the web — written for builders.
Building a content aggregation pipeline with MCP
Map sources, fetch markdown in one agent loop, dedupe by hash, and let monitors push only what changed — a five-tool aggregation pipeline with real calls.
Web scraping best practices for AI training data
Provenance at fetch time, markdown extraction, near-dup removal and deterministic quality filters — the pipeline that turns a raw crawl into training data.
Building a price monitoring system with AI agents
Change detection tells you a page moved; price monitoring needs a number, a timestamp and a rule. The extract schema, cadence math, credit budget and alert design.
Anti-bot detection: what sites do and how modern scrapers handle it
IP reputation, TLS fingerprints, JS challenges, behavioral scoring — the four layers of bot detection, what each measures, and how a scraping API answers each.
Verify an email address before you send
A bad send costs more than a bounce: it costs your domain's reputation. What POST /api/v1/email/verify checks, what each verdict means, and the one rule that keeps a list clean.
Building a real-time data pipeline with a scraping API
Real-time scraping is a latency budget, not a vendor feature. The three cadences, the batch and monitor calls for each, and how to measure end-to-end freshness.
Structured data extraction from e-commerce product pages
JSON-LD, rendered text, or an LLM schema? The three layers of product data, a curl call for each, and the nullable-schema rule that stops silent garbage.
Crawling etiquette: robots.txt, rate limits, and being a good citizen
What robots.txt actually obliges a scraper to do, how to pace a crawl so a small site never notices you, and the identity checks that keep your IP reputation clean.
Why your scraping API bill is unpredictable
Token metering, per-endpoint multipliers and add-on stacking turn a rate card into a variable bill. The four mechanics, and how to price a scrape upfront.
Screenshot APIs for visual regression and archiving
Pixel diffing, deterministic captures and evidence-grade archives — what a screenshot API returns, the settings that make captures comparable, and the cost math.
How to handle pagination in web scraping: offset, cursor, and load-more
Offset, cursor and load-more pagination each break a scraper differently. Find the real endpoint, terminate on evidence, and dedupe rows across page boundaries.
Web scraping for market research: a practical pipeline
Map a competitor set, extract one schema across every source, and let scheduled monitors turn snapshots into a weekly signal — with real curl examples.
Web scraping for SEO monitoring: rankings, on-page, competitors
Rank checks, bulk on-page audits and competitor diffs with real curl examples — plus the cache trap that makes rank data lie, and the credit math.
Scheduled monitors vs polling: change detection that scales
Polling re-fetches everything on your schedule; a monitor re-fetches, hashes and tells you only when something moved. The cost math and the webhook pattern.
API rate limits and credit budgets: plan your scraping capacity
Rate limits cap requests per second; credits cap everything you scrape. The crawl math, 429 backoff that works, and the guardrails that stop runaway crawls.
Extracting structured data from PDFs at scale
PDFs still hold most of the long-form data on the web. The three tiers of PDF extraction, a curl example for each, and the failure modes that kill naive pipelines.
Bright Data vs Fastcrawl: which one should your agent use?
Bright Data is the enterprise proxy-and-data giant; Fastcrawl is a flat rendering API. $1.50/1K records vs flat 1 credit per format, MCP-native, $5/mo.
Decodo vs Fastcrawl: which one should your agent use?
Decodo sells request blocks with a tunable per-1K rate; Fastcrawl is flat. $19 floor vs $5 entry, MCP-native, clean markdown by default.
Apify vs Fastcrawl: which one should your agent use?
Apify is an actor marketplace; Fastcrawl is a flat rendering API. One flat credit per format, MCP-native, $5/mo with 1,500 recurring free credits.
ScrapingBee vs Fastcrawl: 2026 pricing & feature comparison
ScrapingBee's free tier is one-time credits; Fastcrawl's is recurring. $5 entry vs $19, flat credits, MCP-native, clean markdown by default.
Scraping job boards at scale: patterns, pitfalls, and a working pipeline
Map listing URLs, extract title/company/location/salary as JSON, and monitor for fresh postings with webhooks — with real curl examples.
What is a web scraping API? A practical explanation
The plain-language answer to what a scraping API is, why it exists, and when you need one — plus the quick test to tell if yours is good.
How does web scraping work? The 3-step process, explained
Request, extract, save — the always-the-same three steps, and which parts of each step get hard (anti-bot, JS rendering, constant change).
How to convert HTML to Markdown (URL or code)
The practical difference between converting a code block you have and a live web page — and where a library stops being enough.
HTTP-only vs browser rendering: when you don't need a browser
A real browser costs 2-5 seconds per page. Most of the web doesn't need one. The 2-tier fetch_mode pattern (auto, http) and the rule for choosing between them.
How to scrape JavaScript-heavy websites without getting blocked
The three-layer problem — render, extract, evade — and the practical tools that solve each layer for modern SPAs and dynamic sites.
Firecrawl vs Jina Reader vs Fastcrawl: an honest comparison for 2026
Pricing math, latency, and feature surface — what you actually get from each scraping API, and when each one is the right pick.
Cut your LLM bill: stop paying tokens for boilerplate
The biggest cost lever isn't prompts or models — it's what you feed them. Clean extraction, caching, and monitors do most of the work.
What 'LLM-ready markdown' actually means (and how to get it)
Boilerplate stripping, JS rendering, token cost, and structure — the four things that decide whether your RAG pipeline eats good data or garbage.
Give your AI agent the web: connecting an MCP server in one config block
A one-minute setup that turns any MCP-speaking agent into a live-web agent — plus the day-one prompts worth trying.
Monitoring websites with AI agents: change detection that doesn't wake you at 3am
Scheduled re-scrapes, hashing, diffing and webhooks — a practical pattern for tracking competitors, docs, prices and changelogs.
Crawling vs scraping: which one does your project actually need?
Different operations, different costs, different etiquette. The decision rule — plus where 'map' fits in.