Building a content aggregation pipeline with MCP
Most content aggregators are a pile of glue: a script that knows each source's URL shape, a scraper module per site, a cron entry per source, and a diff job bolted on the end. Every new source costs another branch, and the moment a layout moves, the branch silently returns empty. The pipeline itself is not complicated — discover, fetch, normalize, dedupe, deliver — and none of those stages needs per-site code. What it needs is one uniform way to reach the web from wherever your logic already runs.
Model Context Protocol gives you that. A hosted scraping MCP server exposes discovery, fetch, extract and change-detection as plain tools; your agent or orchestrator calls them the same way it calls any other tool, and the per-source knowledge shrinks to a URL and a schema. Everything below is the pipeline I would build first, with the actual config and calls againstfastcrawl.net.
Connect the server once
One config block, and every MCP-speaking client gets the tools: Claude Desktop, Cursor, Codex, Hermes, or your own orchestrator. The endpoint is remote, so there is no package to install and no local process to babysit — the same key you use on the REST API authenticates it.
{
"mcpServers": {
"fastcrawl": {
"url": "https://mcp.fastcrawl.net/mcp",
"headers": { "Authorization": "Bearer ***" }
}
}
}The server ships 14 tools: scrape, crawl,map, search, extract,screenshot, pdf, image,parse, verify_email, and fourmonitor_* tools. For aggregation you will touch five of them. If you have not wired MCP into an agent before, theone-block MCP setup guide walks through the client side in a minute.
The reason to drive this over MCP rather than write REST calls by hand is not aesthetics. Your pipeline logic is a sequence of decisions — which sources to sweep this run, which pages look like the same story, whether a change is worth sending — and an agent already makes that sequence fluently when the tools are in its context. The glue code you would have written becomes the prompt.
Discover sources: map and search, not hardcoded URLs
Aggregation fails first at source discovery. Hardcoded URLs rot, and sitemaps only tell you what the site wants indexed. Two tools cover the discovery phase:map returns the URL surface of a site you already trust, andsearch finds new surfaces you do not know yet.
Map a known source once and keep the result: every article URL on a blog, every release note on a docs site, every entry on a listing page. The map is cheap and the output is just URLs, so re-running it each cycle is how you notice new content without scraping everything. For genuinely new sources, asearch call with include_content fetches the top results' pages in the same call — one round trip instead of search, then a loop of scrapes.
curl -s -X POST https://fastcrawl.net/api/v1/search/ \
-H "Authorization: Bearer ***" \
-H "Content-Type: application/json" \
-d '{"query":"web scraping changelog release notes","limit":5,"include_content":3}'
# -> results[].url + results[].content (markdown, first 3 pages)
# 1 credit per result, 1 credit per bundled page, cache hits freeNote the split: include_content is a REST-only extra; over MCP thesearch tool stays one credit per call, and you fetch pages with a separate scrape. Same credits either way, one call less on the REST side. Pick per client, not per budget.
Fetch and normalize in one pass
The normalize stage is where aggregation pipelines usually rot, because each source brings its own HTML dialect and the code that converts each one grows without bound. Do not convert: fetch markdown. One format, boilerplate already stripped, and every downstream stage — dedupe, summarize, index, embed — reads the same shape regardless of which of forty sources it came from.
Use scrape for the handful of pages that matter this cycle andcrawl when a source needs a breadth-first sweep, thenbatch/scrape on the REST side when you have a list of 50 URLs from a previous map. fetchMode: "http" skips the browser for static and server-rendered pages — about 100ms instead of seconds — and theauto/http rule is what keeps a daily sweep of a hundred sources inside a sane latency budget. If you need fields rather than prose (title, author, date, price), oneextract call with a JSON schema pulls them out of the markdown you already fetched.
Pagination is the one fetch-stage trap that bites aggregators specifically: a "latest posts" page shows ten items, the source has four hundred, and your digest quietly covers the last three days forever. Theoffset, cursor and load-more patterns are worth reading before you trust any listing page.
Dedupe by hash, then judge
Aggregation means many sources describe the same event: a launch covered by twenty outlets, a spec mirrored on three docs sites, your own feed returning a page under a trailing-slash variant. Hash the normalized markdown per document and drop exact collisions first — that alone removes the bulk. Near-duplicates (one outlet rewriting another's paragraph) need shingling, which is a paragraph-level MinHash pass over the text you already have; thededup-before-filter orderapplies here exactly as it does to a training corpus.
Only after dedupe should a model touch the content, and only to rank and summarize the survivors. Clean markdown in means far fewer tokens per document, which is the wholeboilerplate-as-token-cost argument: stripping chrome at fetch time is cheaper than paying for it in every prompt. Store the URL, fetch time and content hash next to each item so a digest entry can be traced back to what you actually read.
Change detection and delivery
Re-fetching every source on every cycle to see whether something moved is the expensive way to run a digest. A monitor_create does the refetch-and-hash on a schedule and tells you only when a page changed, so a cycle that starts with "what's new" begins from a short list instead of a full sweep. Monitors are also how you keep sources you already covered from costing credits every run: the 48-hour cache means a repeat scrape of an unchanged page inside the window is free.
Delivery is your choice and outside the pipeline: a webhook, a digest email, an index your search box queries. Themonitors vs polling mathcovers the cost side — the short version is that polling pays for every source every time, while a monitor pays for a source only when it moved. A weekly digest over 100 sources at 7-day cadence is a few hundred credits a month, not thousands.
The whole loop, in order
Map or search for sources, scrape the changed-or-new pages to markdown, hash and dedupe, extract fields where the digest needs them, summarize the survivors, deliver, and let monitors handle the next cycle's "what's new" question. Five MCP tools, zero per-site code, and the only thing that changes when you add a source is a URL in a list. Themarket-research pipelinefollows the same skeleton with a different schema at the extract stage — which is the point: the skeleton is reusable, the source list is the only variable.
Cost stays flat because every stage is one credit per page, cache hits are free, and failed calls are never charged. Discovery is a credit per result, fetch is a credit per page, extract is a credit per document — you can add up a cycle before you run it. See the credit math for sizing one against your source count and cadence.
Discovery, fetch, extract and monitors are one credit each, cache hits are free, failed calls never charge. Start free · Read the docs