Structured Data Extraction API
Structured data extraction API: a field schema in, records out
POST /extract takes up to 20 URLs and a flat field schema and returns one typed record per page, filled first from what the page itself publishes: schema.org JSON-LD, Open Graph and meta tags. It bills 1 credit per URL fetched successfully; a URL that fails costs nothing.
{ "searchTime": 0.17, "totalResults": 85000, "page": 1, "organic": [ { "position": 1, "title": "31.10. Connection Pools and Data Sources", "link": "www.postgresql.org/docs/…", "domain": "postgresql.org", "snippet": "For an environment without an application se…" }, { "position": 2, "domain": "medium.com", … }, { "position": 3, "domain": "stackoverflow.blog", … } ], "credits": 1}What does a web data extraction API cost on Searlo?
1 credit per URL that was fetched successfully, charged after the work: $0.30 per 1,000 pages at the Scale pack, however many fields you ask for (up to 30) and whether or not a model filled any of them. URLs that fail are not billed.
Before the call runs, your balance has to cover every URL you sent; afterwards only the successful ones are deducted, and the credits field and X-Credits-Deducted header show the exact amount. For comparison, Firecrawl's pricing page adds 4 credits per page for its JSON output format on top of the 1-credit scrape.
- Price
- 1 credit per successful URL: $0.30 per 1,000 at the Scale pack— Searlo pricing
- Per call
- 1–20 URLs, and a schema of 1–30 fields
- Field types
- string, number, integer, float, boolean, array
- Provenance
- Every filled field is tagged json-ld, meta, text or llm
- Free to start
- 3,000 credits, no card— Searlo pricing
- Firecrawl JSON format
- +4 credits per page on top of a 1-credit scrape— firecrawl.dev/pricing
Third-party figures on this page last verified against the sources linked above. Prices change; if you find one of these stale, tell us and we will correct it.
Most pages worth extracting from already describe themselves. A product page usually carries schema.org JSON-LD with its name, price, currency and availability; an article carries its author and publish date; nearly every page has Open Graph tags. Reading those declarations is exact and cheap, and no model improves on a price the site published itself. So /extract reads them first.
You send URLs and a schema: a flat object mapping each field name to a type. Only fields the page's own data leaves empty go to a language model, and only when one is configured and useLlm is on. Every value is converted to the type you declared and records where it came from; anything still empty comes back null and is listed in missingFields.
This page covers extraction only. For a page's prose rather than its fields, the URL to markdown API is the better fit.
How to extract data from a website with the API
Four steps, one of which is billed.
- 01Write a flat schemaMap up to 30 field names to string, number, integer, float, boolean or array. Use names pages tend to use (price, brand, author, datePublished) so the deterministic pass can match them.
- 02Send the URLsPOST /extract with urls (up to 20) or a single url, plus the schema. Set useLlm: false to keep only values the site declared.
- 03Check each rowEvery URL gets a row with success true or false. Failed rows carry an error, are not billed, and can be retried later.
- 04Read the gaps and the sourcesmissingFields lists what the page never stated; _sources says whether each value came from json-ld, meta, text or llm.
Where each value comes from
Four passes, most exact first. A field is filled by the first pass that finds a value for it.
1. The page's JSON-LD
schema.org entities such as Product, Offer, AggregateRating, Article and Organization are flattened and matched to your field names. Tagged json-ld.
2. Open Graph and meta tags
og:title, og:description, og:image, og:price:amount, the canonical link, and author and published-time meta. Tagged meta.
3. Unambiguous text
A price written with $, €, £, USD, EUR or GBP, the first email address, a phone number. Tagged text.
4. Optional model fallback
Only for fields still empty, only when a model is configured and useLlm is true. Told to answer null, never to guess. Tagged llm.
Extract product fields from three pages
The cURL call produces the response shown below it. The Python version keeps only values the sites declared; the Node version flags model-filled fields for review.
curl -X POST "https://api.searlo.tech/api/v1/extract" \
-H "x-api-key: YOUR_API_KEY" \
-H "content-type: application/json" \
-d '{
"urls": [
"https://shop.example.com/p/trail-runner-2",
"https://shop.example.com/p/road-racer",
"https://shop.example.com/p/discontinued"
],
"schema": {
"name": "string",
"brand": "string",
"price": "number",
"currency": "string",
"availability": "string",
"rating": "number",
"reviews": "integer"
}
}'import requests
resp = requests.post(
"https://api.searlo.tech/api/v1/extract",
headers={"x-api-key": "YOUR_API_KEY"},
json={
"urls": [
"https://shop.example.com/p/trail-runner-2",
"https://shop.example.com/p/road-racer",
],
"schema": {
"name": "string",
"price": "number",
"currency": "string",
"availability": "string",
"rating": "number",
},
"useLlm": False, # only values the sites declared; no model inference
},
timeout=180, # URLs in one call are fetched one after another
)
out = resp.json()
print("billed", out["credits"], "credits for", out["totalResults"], "URLs")
for row in out["data"]:
if not row["success"]:
print("not billed, retry later:", row["url"], row["error"])
continue
print(row["url"], row["data"])
print(" sources:", row["_sources"]) # json-ld / meta / text / llm
print(" missing:", row.get("missingFields", [])) # never guessed, always nullconst res = await fetch("https://api.searlo.tech/api/v1/extract", {
method: "POST",
headers: { "x-api-key": "YOUR_API_KEY", "content-type": "application/json" },
body: JSON.stringify({
urls: ["https://shop.example.com/p/trail-runner-2", "https://shop.example.com/p/road-racer"],
schema: { name: "string", brand: "string", price: "number", currency: "string", availability: "string" },
}),
});
const out = await res.json();
console.log("mode:", out.mode, "| model available:", out.llmAvailable, "| credits:", out.credits);
for (const row of out.data.filter((r) => r.success)) {
// Flag values a model inferred, so a person can review them before they ship.
const inferred = Object.entries(row._sources)
.filter(([, source]) => source === "llm")
.map(([field]) => field);
console.log(row.url, row.data, inferred.length ? `review: ${inferred.join(", ")}` : "");
}Response (illustrative)
Field names are the API's own; values are illustrative. Two URLs succeeded and one failed, so credits is 2.
{
"searchParameters": {
"type": "extract",
"urls": 3,
"fields": ["name", "brand", "price", "currency", "availability", "rating", "reviews"]
},
"mode": "structured+llm",
"llmAvailable": true,
"totalResults": 3,
"data": [
{
"url": "https://shop.example.com/p/trail-runner-2",
"success": true,
"data": {
"name": "Trail Runner 2",
"brand": "Example Outdoor",
"price": 129,
"currency": "USD",
"availability": "InStock",
"rating": 4.6,
"reviews": 312
},
"_sources": {
"name": "json-ld",
"brand": "json-ld",
"price": "json-ld",
"currency": "json-ld",
"availability": "json-ld",
"rating": "json-ld",
"reviews": "json-ld"
},
"via": "http",
"llmUsed": false
},
{
"url": "https://shop.example.com/p/road-racer",
"success": true,
"data": {
"name": "Road Racer",
"brand": "Example Outdoor",
"price": 89.5,
"currency": "USD",
"availability": "In stock",
"rating": null,
"reviews": null
},
"_sources": {
"name": "meta",
"brand": "llm",
"price": "text",
"currency": "text",
"availability": "llm"
},
"missingFields": ["rating", "reviews"],
"via": "browser",
"llmUsed": true
},
{
"url": "https://shop.example.com/p/discontinued",
"success": false,
"error": "scrape failed",
"data": null
}
],
"source": "searlo.tech(extract)",
"credits": 2
}Fields the deterministic pass recognises
| Field | Also matches | Filled from |
|---|---|---|
| name | title, headline, productName | JSON-LD name or headline; og:title; the page title |
| description | summary, excerpt, abstract | JSON-LD description; og:description; meta description |
| price | amount, cost | JSON-LD offer price or lowPrice; og:price:amount; a price in the text |
| currency | priceCurrency | JSON-LD offer priceCurrency; og:price:currency; the price's symbol |
| rating | ratingValue, stars, score | JSON-LD aggregateRating ratingValue |
| reviews | reviewCount, ratingCount, numReviews | JSON-LD aggregateRating reviewCount or ratingCount |
| image | imageUrl, thumbnail, photo | JSON-LD image; og:image |
| sku | productId, mpn, gtin | JSON-LD sku, mpn or gtin13 |
| brand | manufacturer, vendor | JSON-LD brand |
| author | byline, writer, creator | JSON-LD author; the author meta tag |
| datePublished | publishedAt, published_date, date | JSON-LD datePublished or dateCreated; article:published_time or a time tag |
| availability | stock, inStock | JSON-LD offer availability, with the schema.org prefix removed |
| url | link, permalink, website | JSON-LD url; the canonical link; og:url |
| emailAddress, contactEmail | JSON-LD email; the first address in the text | |
| phone | telephone, phoneNumber | JSON-LD telephone; a phone number in the text |
| address | location, streetAddress | JSON-LD address, joined into one line |
Schema and request limits
| Rule | Limit |
|---|---|
| URLs per call | 1 to 20 as urls, or a single url |
| Fields per schema | 1 to 30 |
| Field names | Letters, digits and underscores, starting with a letter or underscore, up to 64 characters |
| Field types | string, number, integer, float, boolean, array |
| Nesting | Not supported: the schema is flat |
| useLlm | true by default; false keeps the model out entirely |
| onlyMainContent | true by default; strips navigation, header and footer before extraction |
Extraction cost in dollars
| Pack | Price | Successful URLs | Per 1,000 URLs |
|---|---|---|---|
| Micro | $3.99 | 5,000 | $0.80 |
| Starter | $9.99 | 20,000 | $0.50 |
| Builder | $29.99 | 75,000 | $0.40 |
| Scale | $74.99 | 250,000 | $0.30 |
| Pro | $199.99 | 900,000 | $0.22 |
| Enterprise | $799 | 4,000,000 | $0.20 |
Pricing maths
Monitoring 2,000 product pages once a day is 180,000 credits over 90 days: $54.00 at Scale-pack rates, inside one Scale pack ($74.99 for 250,000 credits). Credits for new accounts are valid for 90 days, so that is the window to size a pack for.
30 fields cost the same as 3, and a page whose gaps a model filled costs the same as one it did not touch. Failed URLs are free, so retrying them costs only what then succeeds.
To find the pages first, POST /map lists a site's URLs for 1 credit per call, and POST /search-and-scrape finds pages by query for 1 credit plus 1 per page scraped.
Searlo vs Firecrawl for structured extraction
Firecrawl adds 4 credits for its JSON format to the 1-credit page scrape, so one structured page is 5 credits (firecrawl.dev/pricing, checked 2026-09-29; the per-1,000 row is 5 times its annual-billing credit rate). Searlo's schema is flat, so for deeply nested output a general LLM extractor may suit you better. Firecrawl and Google are trademarks of their owners; Searlo is independent and affiliated with neither.
| Capability | Searlo | Firecrawl |
|---|---|---|
| Credits for one structured page | 1, and only if the fetch succeeds | 5 (1 for the scrape, 4 for the JSON format) |
| Per 1,000 structured pages | $0.30 at the Scale pack; $0.20–$0.80 by pack | About $3.00 on Scale, $16.00 on Hobby, billed annually |
| Billing model | One-time credit packs | Monthly plans |
| Free credits | 3,000 once, no card | 1,000 every month, no card |
| Firecrawl figures last verified against their published pricing. Prices change; if one is stale, tell us. | ||
Honest limits
/extract is deterministic first on purpose. That makes it cheap and auditable, and it also sets its edges:
- Flat schemas only: no nested objects or arrays of objects. The array type wraps what it finds in a list.
- One record per URL. A category page listing 50 products yields one record, so send the product URLs instead (/map or /crawl can find them).
- Only the 16 field names above have a deterministic source; any other field needs the model fallback or comes back null.
- Text patterns know $, €, £ and USD, EUR, GBP; other currencies need JSON-LD or the model.
- URLs in one call are fetched one after another. For throughput, send several calls in parallel within your rate limit.
- Always a live fetch: there is no cache and no maxAge on /extract.
- Model-filled values are inferences, tagged llm. Review them, or set useLlm: false.
- The model sees a length-capped slice of each page, so a fact deep in a very long page can be missed.
- No logins: pages behind a sign-in cannot be fetched.
What teams extract
- Price and availability monitoring across retailer product pages
- Lead enrichment: name, email, phone and address from company contact pages
- Article datasets: title, author, datePublished and description from news and blog pages
- Catalogue clean-up: sku, brand and image across supplier sites
- Local business records: address, telephone and rating from business pages
- Search, then extract: find pages with /search-and-scrape, then pull fields from the best ones
Where extraction fits
Web scraping API overview
Scrape, map, crawl and extract on one key, with the pricing comparison.
URL to markdown API
When you want the page's prose for a model, not its fields.
Firecrawl alternative
Scrape and crawl pricing next to Firecrawl's published plans.
Google Shopping API
Prices and merchants from Google Shopping results instead of retailer pages.
Search API for RAG
Live search results for retrieval pipelines that also need page fields.
SERP API for LLMs
Structured Google results for grounding, alongside extracted records.
FAQ
Data extraction API FAQ
What is a structured data extraction API?
An API that turns web pages into records with the fields you name. You send URLs and a schema such as name: string, price: number, and get back one record per page with those fields filled, typed, and null wherever the page does not say.
How is /extract billed?
1 credit per URL fetched successfully, charged after the call: $0.30 per 1,000 pages at the Scale pack, whatever the field count and whether or not a model filled some fields. A URL that fails is not billed, though your balance has to cover every URL you send before the call starts.
Is this AI web scraping?
Partly, and only where it has to be. Values the page declares (JSON-LD, Open Graph, meta tags, clear text patterns) are read without a model. A language model fills only the fields still missing, when one is configured and useLlm is true, and each value it fills is tagged llm.
What happens when a field is not on the page?
It comes back null and is listed in missingFields. The model fallback is instructed to return null rather than guess, because a plausible invention is worse than a gap you can see.
Can I extract data from many URLs at once?
Up to 20 per call, fetched one after another. For bigger jobs, run several calls in parallel within your rate limit, or use /map or /crawl to collect the URLs first.
Can the schema be nested?
No. The schema is flat: up to 30 field names, each typed string, number, integer, float, boolean or array. Build nested structures in your own code from the flat record.
Does it work on JavaScript-rendered pages?
Yes. /extract fetches pages the same way /scrape does: a page that arrives as an empty JavaScript shell gets one browser pass automatically, and each row's via field says http or browser.
How is extraction different from scraping a page to markdown?
Scraping returns a page's content for a person or a model to read; extraction returns specific fields for code to use. For an article's text in a RAG index, use /scrape via the URL to markdown API; for price, brand and availability as columns, use /extract.
Extract from your first 3,000 pages free
3,000 free credits, no card. You pay 1 credit per URL that works, and nothing for the ones that fail.