Playbook · 2026-09-28

Structured data extraction from e-commerce product pages

Product pages are the most valuable and least cooperative target on the web. Every retailer wants Google to understand its price, so almost all of them publish some machine-readable version of the product. Almost none of them publish the same one. You get schema.org JSON-LD on some sites, OpenGraph tags on others, microdata on a shrinking minority, and a plain JavaScript-rendered DOM on the rest.

The mistake is treating this as one problem. It is three, with very different costs. Start from the cheapest layer that carries the fields you need and only escalate when it actually comes up empty. This post walks the layers in order, with the real calls and the real output.

Layer 1: ask for the page's own structured data

Fastcrawl's /api/v1/scrape endpoint can return a json format alongside the markdown. That is a deterministic parse of the rendered DOM, not an LLM call: you get the page title, headings, paragraphs, links, images and a word count. It is fast, it is stable, and it costs the same flat credit as any other format.

curl -X POST https://fastcrawl.net/api/v1/scrape \
  -H "Authorization: Bearer ***" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html",
    "formats": ["json", "markdown"]
  }'

The response puts that parse under a json key and the prose under markdown, with metadata and duration_ms alongside:

{
  "success": true,
  "url": "https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html",
  "json": {
    "title": "A Light in the Attic | Books to Scrape - Sandbox",
    "description": "It's hard to imagine a world without A Light in the Attic...",
    "headings": [{ "level": 1, "text": "A Light in the Attic" }, ...],
    "paragraphs": [...],
    "links": [...],
    "images": [...],
    "word_count": 258
  },
  "metadata": { "content_type": "text/html", "duration_ms": 1183 }
}

Use this layer to confirm a URL is a product page at all, to get the exact title string, and to build a cheap change-detection signal. What it will not give you is the price. The DOM parse is deliberately generic, so structured commercial fields like price, currency, SKU and availability are not in it. That is the next layer, and it is where the real work lives.

Layer 2: the schema.org lottery

Open a product page's HTML and search for application/ld+json. When it is there, a Product or ProductGroup block carries name, offers, price, currency and often SKU. That block is the cheapest correct source of truth on the page, because the retailer wrote it for machines.

The lottery is coverage. Measured on a handful of live retail pages: a Shopify storefront (Gymshark) served one ProductGroup JSON-LD block; IKEA's US product page served zero JSON-LD blocks and zero itemprop attributes; a large bookstore served none either. So a pipeline that only reads JSON-LD covers an unpredictable fraction of your target list, and the fraction shifts as retailers change themes.

Read it when it exists, because it is free and exact. Never make it the only path. If you want the HTML to parse yourself, add "html" to formats and pull the block out with a regex on ld\+json. Note the cost: raw HTML for one product page during testing came back at 5.2 MB, which is fine for one page and wasteful for twenty thousand.

Layer 3: a schema, when the page does not cooperate

For the pages with no markup, the reliable route is schema-driven extraction.POST /api/v1/extract fetches the URL, then fills a JSON Schema you supply using the page's content. One call, one row, whatever the markup looked like.

curl -X POST https://fastcrawl.net/api/v1/extract \
  -H "Authorization: Bearer ***" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html",
    "prompt": "Extract the product exactly as listed. Price is the GBP amount shown in the product information table. Availability is the stock text shown.",
    "schema": {
      "type": "object",
      "properties": {
        "title":        { "type": "string" },
        "price":        { "type": "number" },
        "currency":     { "type": "string" },
        "availability": { "type": "string" },
        "upc":          { "type": "string" }
      },
      "required": ["title", "price"]
    }
  }'

Which returns the schema filled in, as a data object:

{
  "success": true,
  "url": ".../a-light-in-the-attic_1000/index.html",
  "data": {
    "title": "A Light in the Attic",
    "price": 51.77,
    "currency": "GBP",
    "availability": "In stock (22 available)",
    "upc": "a897fe39b1053632"
  }
}

Keep the split clean: the prompt carries the disambiguation rules ("the price shown in the product information table", "the currency of the listed price") and theschema carries the structure. When a field is wrong you fix a sentence, and when you need a new field you add a property. Neither change touches the other. This is the same pattern that makes cross-vendor collection work inscraping for market research, just pointed at a catalogue instead of a competitor set.

Make the schema nullable, or you will not know it failed

The single highest-value habit in product extraction is allowing null. Give every optional field a nullable type, and make the prompt say what to do when the page is not a product page. Run the same schema against a category listing and you should get nulls, not invention:

{
  "type": "object",
  "properties": {
    "title":    { "type": ["string", "null"] },
    "price":    { "type": ["number", "null"] },
    "currency": { "type": ["string", "null"] }
  },
  "required": ["title", "price"]
}

Tested against a listing page rather than a product page, that schema returned{"title": null, "price": null, "currency": null} with a 200. That is the correct answer, and it is far more useful to you than a plausible guess. A null is a row you can drop, alert on, or route to a better source. An invented price is a row that silently corrupts your price index and only shows up when someone acts on it.

Two rules follow. Validate required numeric fields before they reach your database, rejecting non-finite and implausible values. And sanity-check the currency of any price before comparing it across retailers, because a schema that allows a string currency in does not guarantee the retailer put the right one in.

Scale it as a batch, not a loop

Product extraction is naturally a list job. POST /api/v1/batch/scrape takes up to 50 URLs and runs them 10-wide, returning a result per URL with its own success flag. A three-URL batch on the test catalogue returned total: 3, succeeded: 3, failed: 0in under a second of engine time. Failures are per item, so one dead listing cannot poison a run.

curl -X POST https://fastcrawl.net/api/v1/batch/scrape \
  -H "Authorization: Bearer ***" \
  -H "Content-Type: application/json" \
  -d '{
    "urls": [
      "https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html",
      "https://books.toscrape.com/catalogue/tipping-the-velvet_999/index.html",
      "https://books.toscrape.com/catalogue/soumission_998/index.html"
    ],
    "formats": ["json"]
  }'

For the URLs that need the third layer, run the batch to learn which ones came back thin, then send only those through /extract. On a catalogue where 70 percent of pages carry JSON-LD, that splits a 20,000-page job into 14,000 cheap parses and 6,000 schema calls instead of 20,000 of the expensive option.

When the page never arrives

Retail pages fail in specific ways, and the failure mode tells you what to do. A short markdown body with no product content means the page is a JavaScript shell; the same request withfetchMode: "auto" escalates to a browser render, andthe HTTP-versus-browser rule explains when that helps. A classified error, like the 422 returned by a bot-walled listing during testing, means the target refused the fetch rather than that your schema was wrong. Do not retry those in a tight loop; fix the fetch path or drop the source.

Watch out for the quiet version too. One large furniture retailer returned 30 KB of markdown for a product URL that contained only navigation and no product fields at all. The request looked successful. Nothing in a naive pipeline would have flagged it. If you are not checking that required fields arrived non-null, you are not measuring success, you are measuring HTTP codes.

Wire the loop up once and product extraction stops being a scraping problem. The layers are the budget: parse what the site already publishes, use a schema for the rest, allow nulls so you can tell the two apart. Endpoint reference and the flat per-call credit rule are in theAPI docs. If your targets are mostly JavaScript-rendered storefronts, readthe JS-heavy scraping guide before you scale this up.

Schema extraction, batches of 50 and deterministic JSON on one flat credit per call. Start free · Read the docs