Structured Data Extraction API

Structured data extraction API: a field schema in, records out

POST /extract takes up to 20 URLs and a flat field schema and returns one typed record per page, filled first from what the page itself publishes: schema.org JSON-LD, Open Graph and meta tags. It bills 1 credit per URL fetched successfully; a URL that fails costs nothing.

GET/api/v1/search/web?q=postgres+connection+pooling&num=3
200 OK0.17s1 credit85,000 results
{  "searchTime": 0.17,  "totalResults": 85000,  "page": 1,  "organic": [    {      "position": 1,      "title": "31.10. Connection Pools and Data Sources",      "link": "www.postgresql.org/docs/…",      "domain": "postgresql.org",      "snippet": "For an environment without an application se…"    },    { "position": 2, "domain": "medium.com", … },    { "position": 3, "domain": "stackoverflow.blog", … }  ],  "credits": 1}

What does a web data extraction API cost on Searlo?

1 credit per URL that was fetched successfully, charged after the work: $0.30 per 1,000 pages at the Scale pack, however many fields you ask for (up to 30) and whether or not a model filled any of them. URLs that fail are not billed.

Before the call runs, your balance has to cover every URL you sent; afterwards only the successful ones are deducted, and the credits field and X-Credits-Deducted header show the exact amount. For comparison, Firecrawl's pricing page adds 4 credits per page for its JSON output format on top of the 1-credit scrape.

Price
1 credit per successful URL: $0.30 per 1,000 at the Scale pack— Searlo pricing
Per call
1–20 URLs, and a schema of 1–30 fields
Field types
string, number, integer, float, boolean, array
Provenance
Every filled field is tagged json-ld, meta, text or llm
Free to start
3,000 credits, no card— Searlo pricing
Firecrawl JSON format
+4 credits per page on top of a 1-credit scrape— firecrawl.dev/pricing

Third-party figures on this page last verified against the sources linked above. Prices change; if you find one of these stale, tell us and we will correct it.

Most pages worth extracting from already describe themselves. A product page usually carries schema.org JSON-LD with its name, price, currency and availability; an article carries its author and publish date; nearly every page has Open Graph tags. Reading those declarations is exact and cheap, and no model improves on a price the site published itself. So /extract reads them first.

You send URLs and a schema: a flat object mapping each field name to a type. Only fields the page's own data leaves empty go to a language model, and only when one is configured and useLlm is on. Every value is converted to the type you declared and records where it came from; anything still empty comes back null and is listed in missingFields.

This page covers extraction only. For a page's prose rather than its fields, the URL to markdown API is the better fit.

How to extract data from a website with the API

Four steps, one of which is billed.

  1. 01Write a flat schemaMap up to 30 field names to string, number, integer, float, boolean or array. Use names pages tend to use (price, brand, author, datePublished) so the deterministic pass can match them.
  2. 02Send the URLsPOST /extract with urls (up to 20) or a single url, plus the schema. Set useLlm: false to keep only values the site declared.
  3. 03Check each rowEvery URL gets a row with success true or false. Failed rows carry an error, are not billed, and can be retried later.
  4. 04Read the gaps and the sourcesmissingFields lists what the page never stated; _sources says whether each value came from json-ld, meta, text or llm.

Where each value comes from

Four passes, most exact first. A field is filled by the first pass that finds a value for it.

  • 1. The page's JSON-LD

    schema.org entities such as Product, Offer, AggregateRating, Article and Organization are flattened and matched to your field names. Tagged json-ld.

  • 2. Open Graph and meta tags

    og:title, og:description, og:image, og:price:amount, the canonical link, and author and published-time meta. Tagged meta.

  • 3. Unambiguous text

    A price written with $, €, £, USD, EUR or GBP, the first email address, a phone number. Tagged text.

  • 4. Optional model fallback

    Only for fields still empty, only when a model is configured and useLlm is true. Told to answer null, never to guess. Tagged llm.

Extract product fields from three pages

The cURL call produces the response shown below it. The Python version keeps only values the sites declared; the Node version flags model-filled fields for review.

curl -X POST "https://api.searlo.tech/api/v1/extract" \
  -H "x-api-key: YOUR_API_KEY" \
  -H "content-type: application/json" \
  -d '{
    "urls": [
      "https://shop.example.com/p/trail-runner-2",
      "https://shop.example.com/p/road-racer",
      "https://shop.example.com/p/discontinued"
    ],
    "schema": {
      "name": "string",
      "brand": "string",
      "price": "number",
      "currency": "string",
      "availability": "string",
      "rating": "number",
      "reviews": "integer"
    }
  }'
import requests

resp = requests.post(
    "https://api.searlo.tech/api/v1/extract",
    headers={"x-api-key": "YOUR_API_KEY"},
    json={
        "urls": [
            "https://shop.example.com/p/trail-runner-2",
            "https://shop.example.com/p/road-racer",
        ],
        "schema": {
            "name": "string",
            "price": "number",
            "currency": "string",
            "availability": "string",
            "rating": "number",
        },
        "useLlm": False,  # only values the sites declared; no model inference
    },
    timeout=180,  # URLs in one call are fetched one after another
)
out = resp.json()
print("billed", out["credits"], "credits for", out["totalResults"], "URLs")

for row in out["data"]:
    if not row["success"]:
        print("not billed, retry later:", row["url"], row["error"])
        continue
    print(row["url"], row["data"])
    print("  sources:", row["_sources"])                 # json-ld / meta / text / llm
    print("  missing:", row.get("missingFields", []))    # never guessed, always null
const res = await fetch("https://api.searlo.tech/api/v1/extract", {
  method: "POST",
  headers: { "x-api-key": "YOUR_API_KEY", "content-type": "application/json" },
  body: JSON.stringify({
    urls: ["https://shop.example.com/p/trail-runner-2", "https://shop.example.com/p/road-racer"],
    schema: { name: "string", brand: "string", price: "number", currency: "string", availability: "string" },
  }),
});
const out = await res.json();
console.log("mode:", out.mode, "| model available:", out.llmAvailable, "| credits:", out.credits);

for (const row of out.data.filter((r) => r.success)) {
  // Flag values a model inferred, so a person can review them before they ship.
  const inferred = Object.entries(row._sources)
    .filter(([, source]) => source === "llm")
    .map(([field]) => field);
  console.log(row.url, row.data, inferred.length ? `review: ${inferred.join(", ")}` : "");
}

Response (illustrative)

Field names are the API's own; values are illustrative. Two URLs succeeded and one failed, so credits is 2.

{
  "searchParameters": {
    "type": "extract",
    "urls": 3,
    "fields": ["name", "brand", "price", "currency", "availability", "rating", "reviews"]
  },
  "mode": "structured+llm",
  "llmAvailable": true,
  "totalResults": 3,
  "data": [
    {
      "url": "https://shop.example.com/p/trail-runner-2",
      "success": true,
      "data": {
        "name": "Trail Runner 2",
        "brand": "Example Outdoor",
        "price": 129,
        "currency": "USD",
        "availability": "InStock",
        "rating": 4.6,
        "reviews": 312
      },
      "_sources": {
        "name": "json-ld",
        "brand": "json-ld",
        "price": "json-ld",
        "currency": "json-ld",
        "availability": "json-ld",
        "rating": "json-ld",
        "reviews": "json-ld"
      },
      "via": "http",
      "llmUsed": false
    },
    {
      "url": "https://shop.example.com/p/road-racer",
      "success": true,
      "data": {
        "name": "Road Racer",
        "brand": "Example Outdoor",
        "price": 89.5,
        "currency": "USD",
        "availability": "In stock",
        "rating": null,
        "reviews": null
      },
      "_sources": {
        "name": "meta",
        "brand": "llm",
        "price": "text",
        "currency": "text",
        "availability": "llm"
      },
      "missingFields": ["rating", "reviews"],
      "via": "browser",
      "llmUsed": true
    },
    {
      "url": "https://shop.example.com/p/discontinued",
      "success": false,
      "error": "scrape failed",
      "data": null
    }
  ],
  "source": "searlo.tech(extract)",
  "credits": 2
}

Fields the deterministic pass recognises

Name fields like these and JSON-LD, meta tags or text patterns can fill them without a model. Matching ignores case, spaces, underscores and hyphens. Any other field name is filled only by the model fallback.
FieldAlso matchesFilled from
nametitle, headline, productNameJSON-LD name or headline; og:title; the page title
descriptionsummary, excerpt, abstractJSON-LD description; og:description; meta description
priceamount, costJSON-LD offer price or lowPrice; og:price:amount; a price in the text
currencypriceCurrencyJSON-LD offer priceCurrency; og:price:currency; the price's symbol
ratingratingValue, stars, scoreJSON-LD aggregateRating ratingValue
reviewsreviewCount, ratingCount, numReviewsJSON-LD aggregateRating reviewCount or ratingCount
imageimageUrl, thumbnail, photoJSON-LD image; og:image
skuproductId, mpn, gtinJSON-LD sku, mpn or gtin13
brandmanufacturer, vendorJSON-LD brand
authorbyline, writer, creatorJSON-LD author; the author meta tag
datePublishedpublishedAt, published_date, dateJSON-LD datePublished or dateCreated; article:published_time or a time tag
availabilitystock, inStockJSON-LD offer availability, with the schema.org prefix removed
urllink, permalink, websiteJSON-LD url; the canonical link; og:url
emailemailAddress, contactEmailJSON-LD email; the first address in the text
phonetelephone, phoneNumberJSON-LD telephone; a phone number in the text
addresslocation, streetAddressJSON-LD address, joined into one line

Schema and request limits

POST /extract body rules, as the API validates them.
RuleLimit
URLs per call1 to 20 as urls, or a single url
Fields per schema1 to 30
Field namesLetters, digits and underscores, starting with a letter or underscore, up to 64 characters
Field typesstring, number, integer, float, boolean, array
NestingNot supported: the schema is flat
useLlmtrue by default; false keeps the model out entirely
onlyMainContenttrue by default; strips navigation, header and footer before extraction

Extraction cost in dollars

1 credit per successful URL, whatever the field count. Pack prices from /pricing; one-time purchases.
PackPriceSuccessful URLsPer 1,000 URLs
Micro$3.995,000$0.80
Starter$9.9920,000$0.50
Builder$29.9975,000$0.40
Scale$74.99250,000$0.30
Pro$199.99900,000$0.22
Enterprise$7994,000,000$0.20

Pricing maths

Monitoring 2,000 product pages once a day is 180,000 credits over 90 days: $54.00 at Scale-pack rates, inside one Scale pack ($74.99 for 250,000 credits). Credits for new accounts are valid for 90 days, so that is the window to size a pack for.

30 fields cost the same as 3, and a page whose gaps a model filled costs the same as one it did not touch. Failed URLs are free, so retrying them costs only what then succeeds.

To find the pages first, POST /map lists a site's URLs for 1 credit per call, and POST /search-and-scrape finds pages by query for 1 credit plus 1 per page scraped.

Searlo vs Firecrawl for structured extraction

Firecrawl adds 4 credits for its JSON format to the 1-credit page scrape, so one structured page is 5 credits (firecrawl.dev/pricing, checked 2026-09-29; the per-1,000 row is 5 times its annual-billing credit rate). Searlo's schema is flat, so for deeply nested output a general LLM extractor may suit you better. Firecrawl and Google are trademarks of their owners; Searlo is independent and affiliated with neither.

Capability comparison between Searlo and Firecrawl
CapabilitySearloFirecrawl
Credits for one structured page1, and only if the fetch succeeds5 (1 for the scrape, 4 for the JSON format)
Per 1,000 structured pages$0.30 at the Scale pack; $0.20–$0.80 by packAbout $3.00 on Scale, $16.00 on Hobby, billed annually
Billing modelOne-time credit packsMonthly plans
Free credits3,000 once, no card1,000 every month, no card
Firecrawl figures last verified against their published pricing. Prices change; if one is stale, tell us.

Honest limits

/extract is deterministic first on purpose. That makes it cheap and auditable, and it also sets its edges:

  • Flat schemas only: no nested objects or arrays of objects. The array type wraps what it finds in a list.
  • One record per URL. A category page listing 50 products yields one record, so send the product URLs instead (/map or /crawl can find them).
  • Only the 16 field names above have a deterministic source; any other field needs the model fallback or comes back null.
  • Text patterns know $, €, £ and USD, EUR, GBP; other currencies need JSON-LD or the model.
  • URLs in one call are fetched one after another. For throughput, send several calls in parallel within your rate limit.
  • Always a live fetch: there is no cache and no maxAge on /extract.
  • Model-filled values are inferences, tagged llm. Review them, or set useLlm: false.
  • The model sees a length-capped slice of each page, so a fact deep in a very long page can be missed.
  • No logins: pages behind a sign-in cannot be fetched.

What teams extract

  • Price and availability monitoring across retailer product pages
  • Lead enrichment: name, email, phone and address from company contact pages
  • Article datasets: title, author, datePublished and description from news and blog pages
  • Catalogue clean-up: sku, brand and image across supplier sites
  • Local business records: address, telephone and rating from business pages
  • Search, then extract: find pages with /search-and-scrape, then pull fields from the best ones

Where extraction fits

FAQ

Data extraction API FAQ

What is a structured data extraction API?

An API that turns web pages into records with the fields you name. You send URLs and a schema such as name: string, price: number, and get back one record per page with those fields filled, typed, and null wherever the page does not say.

How is /extract billed?

1 credit per URL fetched successfully, charged after the call: $0.30 per 1,000 pages at the Scale pack, whatever the field count and whether or not a model filled some fields. A URL that fails is not billed, though your balance has to cover every URL you send before the call starts.

Is this AI web scraping?

Partly, and only where it has to be. Values the page declares (JSON-LD, Open Graph, meta tags, clear text patterns) are read without a model. A language model fills only the fields still missing, when one is configured and useLlm is true, and each value it fills is tagged llm.

What happens when a field is not on the page?

It comes back null and is listed in missingFields. The model fallback is instructed to return null rather than guess, because a plausible invention is worse than a gap you can see.

Can I extract data from many URLs at once?

Up to 20 per call, fetched one after another. For bigger jobs, run several calls in parallel within your rate limit, or use /map or /crawl to collect the URLs first.

Can the schema be nested?

No. The schema is flat: up to 30 field names, each typed string, number, integer, float, boolean or array. Build nested structures in your own code from the flat record.

Does it work on JavaScript-rendered pages?

Yes. /extract fetches pages the same way /scrape does: a page that arrives as an empty JavaScript shell gets one browser pass automatically, and each row's via field says http or browser.

How is extraction different from scraping a page to markdown?

Scraping returns a page's content for a person or a model to read; extraction returns specific fields for code to use. For an article's text in a RAG index, use /scrape via the URL to markdown API; for price, brand and availability as columns, use /extract.

Extract from your first 3,000 pages free

3,000 free credits, no card. You pay 1 credit per URL that works, and nothing for the ones that fail.