Notes on web data.

How AI agents read, extract and monitor the web — written for builders.

Playbook · 2026-10-09

Building a content aggregation pipeline with MCP

Map sources, fetch markdown in one agent loop, dedupe by hash, and let monitors push only what changed — a five-tool aggregation pipeline with real calls.

Playbook · 2026-10-07

Web scraping best practices for AI training data

Provenance at fetch time, markdown extraction, near-dup removal and deterministic quality filters — the pipeline that turns a raw crawl into training data.

Playbook · 2026-10-05

Building a price monitoring system with AI agents

Change detection tells you a page moved; price monitoring needs a number, a timestamp and a rule. The extract schema, cadence math, credit budget and alert design.

Fundamentals · 2026-10-02

Anti-bot detection: what sites do and how modern scrapers handle it

IP reputation, TLS fingerprints, JS challenges, behavioral scoring — the four layers of bot detection, what each measures, and how a scraping API answers each.

Playbook · 2026-10-01

Verify an email address before you send

A bad send costs more than a bounce: it costs your domain's reputation. What POST /api/v1/email/verify checks, what each verdict means, and the one rule that keeps a list clean.

Playbook · 2026-09-30

Building a real-time data pipeline with a scraping API

Real-time scraping is a latency budget, not a vendor feature. The three cadences, the batch and monitor calls for each, and how to measure end-to-end freshness.

Playbook · 2026-09-28

Structured data extraction from e-commerce product pages

JSON-LD, rendered text, or an LLM schema? The three layers of product data, a curl call for each, and the nullable-schema rule that stops silent garbage.

Fundamentals · 2026-09-25

Crawling etiquette: robots.txt, rate limits, and being a good citizen

What robots.txt actually obliges a scraper to do, how to pace a crawl so a small site never notices you, and the identity checks that keep your IP reputation clean.

Cost engineering · 2026-09-23

Why your scraping API bill is unpredictable

Token metering, per-endpoint multipliers and add-on stacking turn a rate card into a variable bill. The four mechanics, and how to price a scrape upfront.

Playbook · 2026-09-21

Screenshot APIs for visual regression and archiving

Pixel diffing, deterministic captures and evidence-grade archives — what a screenshot API returns, the settings that make captures comparable, and the cost math.

Playbook · 2026-09-18

How to handle pagination in web scraping: offset, cursor, and load-more

Offset, cursor and load-more pagination each break a scraper differently. Find the real endpoint, terminate on evidence, and dedupe rows across page boundaries.

Playbook · 2026-09-16

Web scraping for market research: a practical pipeline

Map a competitor set, extract one schema across every source, and let scheduled monitors turn snapshots into a weekly signal — with real curl examples.

Playbook · 2026-09-14

Web scraping for SEO monitoring: rankings, on-page, competitors

Rank checks, bulk on-page audits and competitor diffs with real curl examples — plus the cache trap that makes rank data lie, and the credit math.

Playbook · 2026-09-11

Scheduled monitors vs polling: change detection that scales

Polling re-fetches everything on your schedule; a monitor re-fetches, hashes and tells you only when something moved. The cost math and the webhook pattern.

Cost engineering · 2026-09-11

API rate limits and credit budgets: plan your scraping capacity

Rate limits cap requests per second; credits cap everything you scrape. The crawl math, 429 backoff that works, and the guardrails that stop runaway crawls.

Playbook · 2026-09-04

Extracting structured data from PDFs at scale

PDFs still hold most of the long-form data on the web. The three tiers of PDF extraction, a curl example for each, and the failure modes that kill naive pipelines.

Comparison · 2026-09-02

Bright Data vs Fastcrawl: which one should your agent use?

Bright Data is the enterprise proxy-and-data giant; Fastcrawl is a flat rendering API. $1.50/1K records vs flat 1 credit per format, MCP-native, $5/mo.

Comparison · 2026-09-02

Decodo vs Fastcrawl: which one should your agent use?

Decodo sells request blocks with a tunable per-1K rate; Fastcrawl is flat. $19 floor vs $5 entry, MCP-native, clean markdown by default.

Comparison · 2026-09-02

Apify vs Fastcrawl: which one should your agent use?

Apify is an actor marketplace; Fastcrawl is a flat rendering API. One flat credit per format, MCP-native, $5/mo with 1,500 recurring free credits.

Comparison · 2026-09-02

ScrapingBee vs Fastcrawl: 2026 pricing & feature comparison

ScrapingBee's free tier is one-time credits; Fastcrawl's is recurring. $5 entry vs $19, flat credits, MCP-native, clean markdown by default.

Playbook · 2026-09-02

Scraping job boards at scale: patterns, pitfalls, and a working pipeline

Map listing URLs, extract title/company/location/salary as JSON, and monitor for fresh postings with webhooks — with real curl examples.

Fundamentals · 2026-09-01

What is a web scraping API? A practical explanation

The plain-language answer to what a scraping API is, why it exists, and when you need one — plus the quick test to tell if yours is good.

Fundamentals · 2026-09-01

How does web scraping work? The 3-step process, explained

Request, extract, save — the always-the-same three steps, and which parts of each step get hard (anti-bot, JS rendering, constant change).

How-to · 2026-09-01

How to convert HTML to Markdown (URL or code)

The practical difference between converting a code block you have and a live web page — and where a library stops being enough.

Fundamentals · 2026-08-31

HTTP-only vs browser rendering: when you don't need a browser

A real browser costs 2-5 seconds per page. Most of the web doesn't need one. The 2-tier fetch_mode pattern (auto, http) and the rule for choosing between them.

How-to · 2026-08-27

How to scrape JavaScript-heavy websites without getting blocked

The three-layer problem — render, extract, evade — and the practical tools that solve each layer for modern SPAs and dynamic sites.

Comparison · 2026-08-25

Firecrawl vs Jina Reader vs Fastcrawl: an honest comparison for 2026

Pricing math, latency, and feature surface — what you actually get from each scraping API, and when each one is the right pick.

Cost engineering · 2026-08-25

Cut your LLM bill: stop paying tokens for boilerplate

The biggest cost lever isn't prompts or models — it's what you feed them. Clean extraction, caching, and monitors do most of the work.

Fundamentals · 2026-08-25

What 'LLM-ready markdown' actually means (and how to get it)

Boilerplate stripping, JS rendering, token cost, and structure — the four things that decide whether your RAG pipeline eats good data or garbage.

Playbook · 2026-08-25

Give your AI agent the web: connecting an MCP server in one config block

A one-minute setup that turns any MCP-speaking agent into a live-web agent — plus the day-one prompts worth trying.

Playbook · 2026-08-25

Monitoring websites with AI agents: change detection that doesn't wake you at 3am

Scheduled re-scrapes, hashing, diffing and webhooks — a practical pattern for tracking competitors, docs, prices and changelogs.

Fundamentals · 2026-08-25

Crawling vs scraping: which one does your project actually need?

Different operations, different costs, different etiquette. The decision rule — plus where 'map' fits in.