Playbook · 2026-10-05

Building a price monitoring system with AI agents

A price monitor has one job: say what this product costs right now, and how that compares with the last time you looked. Most setups in the wild answer a narrower question, namely whether the page changed. The gap between those two questions is where price systems go wrong, and it is not a small gap: one produces a trend line you can query, the other produces a pager that goes off because someone rotated a banner.

This is the whole system, end to end: decide what a price is, pull it out of the page as a number, store it with a timestamp, refresh it at a cadence you can actually afford, and alert on rules that survive contact with a real catalog. Every call below is a real request against fastcrawl.net, and the credit math is stated as you go.

Change detection is not price monitoring

A hash monitor normalizes a page to markdown, hashes it, and fires a webhook when the hash moves. That is the right primitive for a changelog, a docs page or a competitor's pricing page, and it is the wrong primitive for a catalog. A new review count, a swapped hero image, an "only 3 left" badge and an A/B test all move the hash without moving the price. When the price does move, the hash hands you no number, no direction and no magnitude.

The hash-first pipeline therefore re-fetches the page after every alert to find out what actually changed: a second credit per alert, plus a race against a page that has moved again. The fix is to make the value the primitive instead. One row per capture, keyed by retailer and product id, carrying price, currency, stock state and capture time. The comparison lives in your table, where you can query it, backfill it and explain a jump three weeks later.

Getting the value out of the page is its own problem, and the three layers of product data extraction cover it in depth: JSON-LD when the retailer ships it, a schema-driven extract when they do not. The rest of this post assumes you have that layer and builds the system around it.

Get the number out of the page

For a system, use one path for every retailer: a fixed JSON Schema and a schema-driven extract. One call, one row, identical keys whether the target is Shopify, an IKEA-style theme or a page with no structured markup at all. POST /api/v1/extract fetches the URL, then fills your schema from the page content.

curl -s -X POST https://fastcrawl.net/api/v1/extract \
  -H "Authorization: Bearer ***" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://shop.example.com/products/widget-pro",
    "prompt": "Return the price a buyer pays today for the default variant. Use null when no price is shown. Do not return a struck-through compare-at price unless that is the price charged.",
    "schema": {
      "type": "object",
      "properties": {
        "currency":   { "type": "string" },
        "price":      { "type": ["number", "null"] },
        "list_price": { "type": ["number", "null"] },
        "in_stock":   { "type": "boolean" },
        "variant":    { "type": ["string", "null"] }
      },
      "required": ["currency", "price", "in_stock"]
    }
  }'
# -> {"success":true,"url":"...","data":{"currency":"USD","price":49.0,...}}

Two rules keep this honest. First, every field that can legitimately vanish is nullable, and the prompt says what to do when it does. A missing price becomes null, which is a signal you can alert on, instead of a plausible number the model invented to satisfy the schema. Second, use the same schema everywhere and let the prompt vary per domain. One extract call is one credit, failed calls are never charged, and a null row costs the same as a good one.

Do not be tempted to read the price out of markdown with a regex per retailer. It works for the first three targets, then a locale formats 1.299,00 differently and your history has a silent hole where the numbers used to be.

Normalize before you store

A number is only comparable if the units match, and retail pages obscure units for a living. Currency is the obvious one: store the raw price and the currency, and convert in the query, never at ingest. Variant is the second: "default variant" and "cheapest variant" are different products with different prices, so capture the variant string alongside the number and your diffs become explainable.

Then the traps. Struck-through compare-at prices belong in list_price, not in price. Coupon and app-only pricing means the visible number is not the charged number, so the prompt should ask for what a buyer pays. A "from $19" listing page is not a product price at all, and neither is a price shown only after a membership login, which you usually cannot see without one.

Store twice: an append-only history of (retailer, product_id, captured_at, price, currency, in_stock, variant), and a current-row table upserted on (retailer, product_id). Key on the product id the page exposes, not on the URL, because campaign parameters and session paths change while the SKU does not. Keep the source URL in the row for the day you need to re-open the page and argue with the diff.

Cadence and the credit math

Monitors run on a daily, weekly or monthly schedule at an hour you pick, in your timezone. That is the right shape for the standing watch: one capture per SKU per day, one credit per run, and a webhook that fires only when the normalized content actually moved. Monitors and webhooks are included on the $5 Go plan.

curl -s -X POST https://fastcrawl.net/api/v1/monitors \
  -H "Authorization: Bearer ***" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://shop.example.com/products/widget-pro",
    "schedule": "daily",
    "timezone": "Asia/Hong_Kong",
    "hour": 9,
    "webhook_url": "https://api.example.com/hooks/prices"
  }'
# -> {"success":true,"id":"mon_..."}  webhook: monitor.run.completed
#    data = {monitor_id, url, changed, content_hash, markdown (first 50k chars)}

For anything faster than a daily check, you own the schedule: sweep the watchlist with POST /api/v1/batch/scrape, which takes up to 50 URLs per call, runs them 10 wide in parallel, and costs one credit per URL. Set maxAge to your sweep interval so each window pays for at most one fetch and the rest are cache hits, which are free. The credit and capacity math has the backpressure details; the short version is that failed calls never charge, so a retry loop that respects the retryable flag costs nothing extra.

The budget at 500 SKUs on a daily cadence is 500 credits a day, about 15,000 a month. On the Go plan that is 5,000 included plus 10,000 overage at $1.00 per 1,000, so roughly $15 a month for 15,000 fresh captures. The free tier's 1,500 credits cover the same sweep for three days, which is enough to prove the pipeline before you pay for it. Sub-daily sweeps scale that number linearly, which is the real reason to check volatile SKUs hourly and everything else daily.

Where the AI agent earns its place

Fastcrawl exposes the same surface over MCP as over REST, 13 tools on one key. Point any MCP-speaking agent at it and the watchlist becomes something you can drive in plain language: find the product pages, extract the schema, create the monitors, read the runs.

{
  "mcpServers": {
    "fastcrawl": {
      "url": "https://mcp.fastcrawl.net/mcp",
      "headers": { "Authorization": "Bearer YOUR_API_KEY" }
    }
  }
}

Give the agent judgment tasks, not measurement tasks. It is good at triaging a webhook stream (is this a price move or a template edit?), drafting the alert text, deciding which twenty SKUs deserve a re-check after a retailer-wide drop, and proposing threshold changes when a category turns volatile. It is bad at producing the price. An agent reading markdown will occasionally return a compare-at price or a monthly installment as the price, and that bug will look like data for months. The number comes from the schema call; the agent interprets what the number means.

The wiring is boring on purpose: webhook to your endpoint, hand the agent the markdown and hash from the payload, let it call extract for the structured row and decide whether a human should see this. If it decides no, the row is still written. Only the alert is suppressed, so the store stays the ground truth and the agent stays a filter in front of it.

Alert rules that survive a week

Absolute thresholds mislead across categories: $5 off a $30 item and $5 off a $900 item are not the same event. Work in percentages, and add hysteresis so a price oscillating around the boundary does not flap. A practical pair is "alert on a drop of 5% or more, re-arm only after it recovers 3%", which turns a jittering price into one event instead of eleven.

Deduplicate on (product_id, price) within a 24-hour window, and digest anything catalog-wide instead of firing per item: a retailer-wide sale should arrive as one message listing the biggest movers, not four hundred. Before a move reaches a person, confirm it with one fresh fetch (maxAge: 0) so a cached row, a geo variant or an A/B test never turns into a notification. If the confirmation fails, send nothing; failed calls are free, so the check costs nothing either way.

That is the whole system: a nullable schema for the number, a history table for the comparison, daily monitors plus a batched sweep for the cadence, and an agent that filters noise instead of guessing values. Everything above runs against the free tier first, and the pieces are individually boring, which is the point. Price monitoring fails on cleverness, not on missing features.

Schema extracts, batched sweeps and scheduled monitors are one credit each, and failed calls are free. Start free · Read the docs