Web Crawler API
Web Crawler API: crawl a whole site, pay per page fetched
Searlo's web crawler API crawls an entire website as a background job and hands back every page as markdown, text, HTML or links. POST /crawl costs 1 credit per page it actually fetches, and POST /map lists the site's URLs for 1 credit first, so you know the size of the job before you pay for it.
{ "searchTime": 0.17, "totalResults": 85000, "page": 1, "organic": [ { "position": 1, "title": "31.10. Connection Pools and Data Sources", "link": "www.postgresql.org/docs/…", "domain": "postgresql.org", "snippet": "For an environment without an application se…" }, { "position": 2, "domain": "medium.com", … }, { "position": 3, "domain": "stackoverflow.blog", … } ], "credits": 1}How much does it cost to crawl a website with an API?
On Searlo, POST /crawl costs 1 credit per page it successfully fetches, charged while the job runs. At the Scale pack's $0.30 per 1,000 credits, a 500-page crawl costs $0.15 and a job at the 5,000-page maximum costs $1.50. The limit you set is a ceiling, not a quote: a crawl capped at 500 that finds 40 pages is billed 40 credits.
Pages that fail to load and pages robots.txt disallows are not billed. POST /map lists up to 5,000 of a site's URLs from its sitemaps for 1 credit per call, or 2 when you pass maxAge or cache:false and the answer has to be fetched live. A crawl needs at least 5 credits in the account to start.
- Crawl
- 1 credit per page crawled: $0.30 per 1,000 pages at the Scale pack— Searlo pricing
- Map
- 1 credit per call; 2 if maxAge or cache:false forces a live fetch
- Per job
- Up to 5,000 pages, depth 0–10, 1–8 pages fetched in parallel
- Status, results, cancel
- Free: each page is billed once, when it is fetched
- Free to start
- 3,000 credits, no card— Searlo pricing
- Firecrawl basic crawl
- $0.599 to $3.20 per 1,000 pages by plan, billed annually— firecrawl.dev/pricing
Third-party figures on this page last verified against the sources linked above. Prices change; if you find one of these stale, tell us and we will correct it.
A crawl is a job, not a request. You send a start URL and a few rules (how deep to go, which paths to keep, how many pages at most), and POST /crawl answers at once with HTTP 202 and a job id. Workers then read robots.txt, seed the queue from the site's sitemaps, fetch each page through Searlo's managed proxy pool, follow in-scope links and store every page as it completes. You poll the job or give it a webhook, then read the pages back in cursor-paginated batches.
Billing follows the work. Credits are taken in small batches while the job runs, so creditsCharged tracks what was actually fetched. If your balance runs out mid-crawl, the job stops with stopReason "credit balance exhausted", keeps every page it already stored, and charges nothing for pages it never reached.
This page covers crawling a whole site. For one known page, or for fields rather than page bodies, the smaller tools below fit better.
Where crawling fits
Crawling is one of five web-data calls that share one API key and one credit balance.
Web scraping API overview
How scrape, map, crawl and extract fit together.
URL to markdown API
One known page as markdown with token-sized chunks.
Data extraction API
Crawled URLs in, typed records out.
Firecrawl alternative
Scrape and crawl rates next to Firecrawl's plans.
Search API for RAG
Live search results beside a crawled knowledge base.
SERP API for LLMs
Google results a model can cite next to your pages.
How to crawl a website with the API
Five calls. Only the first two can cost anything.
- 01Map the site firstPOST /map returns the URLs the site's robots.txt and sitemaps declare (1,000 by default, up to 5,000 with limit), each with lastModified when the sitemap gives one. It costs 1 credit and sizes the crawl.
- 02Start the crawlPOST /crawl with url, a limit, maxDepth and any path globs. The answer is HTTP 202 with a jobId, statusUrl, resultsUrl and maxCredits, the most the job can cost.
- 03Wait for it to finishPoll GET /crawl/{jobId} for status and progress (discovered, completed, failed, skipped), or pass a webhook for one signed POST when the job ends.
- 04Page through the resultsGET /crawl/{jobId}/results returns up to 200 pages per call with nextCursor and hasMore. Filter by status to see failed or skipped pages and why.
- 05Cancel if you need toDELETE /crawl/{jobId} stops a running crawl. Pages already fetched stay readable, and nothing more is charged.
Map, crawl and collect
The same flow in each language: size the site, start the job, wait, then page through what it stored. Replace YOUR_API_KEY with a key from the dashboard.
# 1. See what the site publishes before you pay to crawl it (1 credit)
curl -X POST "https://api.searlo.tech/api/v1/map" \
-H "x-api-key: YOUR_API_KEY" \
-H "content-type: application/json" \
-d '{ "url": "https://docs.example.com", "search": "/guides/", "limit": 5000 }'
# 2. Start the crawl: 202 Accepted with a job id, billed per page as it runs
curl -X POST "https://api.searlo.tech/api/v1/crawl" \
-H "x-api-key: YOUR_API_KEY" \
-H "content-type: application/json" \
-d '{
"url": "https://docs.example.com/guides/",
"limit": 500,
"maxDepth": 3,
"includePaths": ["/guides/**"],
"excludePaths": ["/guides/archive/**"],
"formats": ["markdown", "links"],
"webhook": { "url": "https://your-app.example.com/hooks/crawl", "secret": "YOUR_WEBHOOK_SECRET" }
}'
# 3. Check progress (free)
curl "https://api.searlo.tech/api/v1/crawl/crawl_5f2c9e0b7a41d3e8c6b1f0a2" \
-H "x-api-key: YOUR_API_KEY"
# 4. Page through the stored pages, up to 200 per call (free)
curl "https://api.searlo.tech/api/v1/crawl/crawl_5f2c9e0b7a41d3e8c6b1f0a2/results?limit=200" \
-H "x-api-key: YOUR_API_KEY"import time
import requests
API = "https://api.searlo.tech/api/v1"
HEADERS = {"x-api-key": "YOUR_API_KEY"}
# 1. Size the job before paying for it (1 credit).
site = requests.post(f"{API}/map", headers=HEADERS,
json={"url": "https://docs.example.com", "limit": 5000}).json()
print(site["totalResults"], "URLs found via", site["discovery"])
# 2. Start the crawl. Credits are charged per page fetched, not per page allowed.
job = requests.post(f"{API}/crawl", headers=HEADERS, json={
"url": "https://docs.example.com/guides/",
"limit": 500,
"maxDepth": 3,
"includePaths": ["/guides/**"],
"formats": ["markdown"],
}).json()
job_id = job["jobId"]
# 3. Poll until the job stops. Status reads are free.
while True:
status = requests.get(f"{API}/crawl/{job_id}", headers=HEADERS).json()
if status["status"] in ("completed", "failed", "cancelled"):
break
time.sleep(10)
print(status["progress"], status["creditsCharged"], status["stopReason"])
# 4. Read every stored page, 200 at a time. Also free: pages were billed when fetched.
cursor = None
while True:
params = {"limit": 200, **({"cursor": cursor} if cursor else {})}
batch = requests.get(f"{API}/crawl/{job_id}/results",
headers=HEADERS, params=params).json()
for page in batch["pages"]:
print(page["url"], page["statusCode"], page["wordCount"])
if not batch["hasMore"]:
break
cursor = batch["nextCursor"]const API = "https://api.searlo.tech/api/v1";
const headers = { "x-api-key": "YOUR_API_KEY", "content-type": "application/json" };
const sleep = (ms) => new Promise((resolve) => setTimeout(resolve, ms));
// Start the crawl: HTTP 202 with a job id.
const job = await fetch(`${API}/crawl`, {
method: "POST",
headers,
body: JSON.stringify({
url: "https://docs.example.com/guides/",
limit: 500,
maxDepth: 3,
includePaths: ["/guides/**"],
formats: ["markdown"],
}),
}).then((r) => r.json());
// Poll the job (free) until it stops.
let status;
do {
await sleep(10_000);
status = await fetch(`${API}/crawl/${job.jobId}`, { headers }).then((r) => r.json());
} while (status.status === "queued" || status.status === "running");
console.log(status.progress, status.creditsCharged, status.stopReason);
// Page through the stored results with the cursor (free).
let cursor = null;
do {
const qs = new URLSearchParams({ limit: "200", ...(cursor && { cursor }) });
const batch = await fetch(`${API}/crawl/${job.jobId}/results?${qs}`, { headers })
.then((r) => r.json());
for (const page of batch.pages) console.log(page.url, page.title, page.wordCount);
cursor = batch.hasMore ? batch.nextCursor : null;
} while (cursor);What the crawler returns (illustrative)
Field names are the API's own; values are illustrative and the result page is trimmed to one entry. The request also accepts snake_case aliases such as max_depth and include_paths.
{
"success": true,
"jobId": "crawl_5f2c9e0b7a41d3e8c6b1f0a2",
"status": "queued",
"url": "https://docs.example.com/guides",
"options": {
"maxDepth": 3,
"limit": 500,
"includePaths": ["/guides/**"],
"excludePaths": ["/guides/archive/**"],
"formats": ["markdown", "links"],
"allowSubdomains": false,
"allowExternal": false,
"respectRobots": true,
"concurrency": 3,
"render": null,
"onlyMainContent": true,
"useSitemap": true,
"storeContent": true
},
"maxCredits": 500,
"statusUrl": "/api/v1/crawl/crawl_5f2c9e0b7a41d3e8c6b1f0a2",
"resultsUrl": "/api/v1/crawl/crawl_5f2c9e0b7a41d3e8c6b1f0a2/results",
"createdAt": "2026-09-29T09:00:00.000Z"
}{
"success": true,
"jobId": "crawl_5f2c9e0b7a41d3e8c6b1f0a2",
"status": "completed",
"url": "https://docs.example.com/guides",
"options": { "maxDepth": 3, "limit": 500, "formats": ["markdown", "links"] },
"progress": { "discovered": 212, "completed": 187, "failed": 3, "skipped": 22 },
"creditsCharged": 187,
"stopReason": null,
"webhook": {
"url": "https://your-app.example.com/hooks/crawl",
"deliveredAt": "2026-09-29T09:06:41.000Z",
"attempts": 1
},
"createdAt": "2026-09-29T09:00:00.000Z",
"startedAt": "2026-09-29T09:00:02.000Z",
"finishedAt": "2026-09-29T09:06:40.000Z"
}{
"success": true,
"jobId": "crawl_5f2c9e0b7a41d3e8c6b1f0a2",
"status": "completed",
"totalResults": 1,
"hasMore": true,
"nextCursor": "66f91a2b3c4d5e6f7a8b9c0d",
"pages": [
{
"url": "https://docs.example.com/guides/install",
"finalUrl": null,
"depth": 1,
"parentUrl": "https://docs.example.com/guides",
"status": "completed",
"statusCode": 200,
"via": "http",
"title": "Install the CLI",
"description": "Install and configure the command-line tool.",
"language": "en",
"canonicalUrl": "https://docs.example.com/guides/install",
"wordCount": 842,
"markdown": "# Install the CLI\n\nRun the installer for your platform, then sign in...",
"links": ["https://docs.example.com/guides/configure"],
"fetchedAt": "2026-09-29T09:02:11.000Z"
}
]
}{
"searchParameters": { "url": "https://docs.example.com/", "type": "map", "limit": 5000, "search": "/guides/" },
"origin": "https://docs.example.com/",
"discovery": "sitemap",
"sitemaps": ["https://docs.example.com/sitemap.xml"],
"robots": { "present": true, "sitemapCount": 1 },
"totalResults": 2,
"urls": [
{ "url": "https://docs.example.com/guides/install", "lastModified": "2026-09-12" },
{ "url": "https://docs.example.com/guides/configure", "lastModified": "2026-08-30" }
],
"source": "searlo.tech(map)",
"credits": 1
}Crawl parameters
| Parameter | Default | Allowed values | What it does |
|---|---|---|---|
| url | required | http(s), up to 2,048 characters | Where the crawl starts; it stays on this host unless you widen it |
| limit | 100 | 1–5,000 | Most pages the job will fetch: a ceiling on cost, not a quote |
| maxDepth | 2 | 0–10 | Link hops from the start URL; sitemap URLs enter at depth 1 |
| includePaths / excludePaths | none | Up to 25 globs each | * matches within a path segment, ** across segments, a plain /docs is a prefix |
| formats | markdown | markdown, text, html, links, json_ld | What is stored for each page |
| concurrency | 3 | 1–8 | Pages fetched in parallel inside the job |
| respectRobots | true | true / false | Skip paths robots.txt disallows; skipped pages are not billed |
| useSitemap | true | true / false | Seed the queue from the site's sitemaps |
| allowSubdomains / allowExternal | false | true / false | Follow links to subdomains, or to other sites entirely |
| render | automatic | true / false | Force a browser render for every page, or never use one |
| onlyMainContent | true | true / false | Strip navigation, header and footer before storing |
| storeContent | true | true / false | false stores each page's title, status and word count (plus links if requested), not its body |
| webhook | none | { url, secret } | One POST when the job ends, HMAC-signed when you set a secret |
Crawl cost in dollars
| Pack | Price | Credits (= pages) | Full 5,000-page jobs | Per 1,000 pages |
|---|---|---|---|---|
| Micro | $3.99 | 5,000 | 1 | $0.80 |
| Starter | $9.99 | 20,000 | 4 | $0.50 |
| Builder | $29.99 | 75,000 | 15 | $0.40 |
| Scale | $74.99 | 250,000 | 50 | $0.30 |
| Pro | $199.99 | 900,000 | 180 | $0.22 |
| Enterprise | $799 | 4,000,000 | 800 | $0.20 |
Worked examples
A 40-page documentation site crawled with limit 500 costs 40 credits, $0.012 at Scale-pack rates; mapping it first adds 1 credit.
A weekly re-crawl of a 2,000-page blog is 26,000 credits over 13 weeks: $7.80 at Scale-pack rates, well inside one Builder pack ($29.99 for 75,000 credits). Credits for new accounts are valid for 90 days, so size a pack to that window.
100,000 pages across many jobs is $30.00 at Scale-pack rates, or $43.97 for exactly 100,000 credits (Builder, Starter and Micro packs together).
Searlo vs Firecrawl for crawling
Both price a basic crawl at one credit per page; the difference is what a credit costs. Firecrawl's figures are from firecrawl.dev/pricing, checked 2026-09-29. Firecrawl and Google are trademarks of their owners; Searlo is independent and affiliated with neither.
| Capability | Searlo | Firecrawl |
|---|---|---|
| Price per 1,000 pages crawled | $0.30 at the Scale pack; $0.20–$0.80 by pack | $0.599 (Scale) to $3.20 (Hobby), billed annually |
| 100,000 pages | $43.97 one-time for exactly 100,000 credits | $83 a month for the Standard plan, billed annually |
| Billing model | One-time credit packs | Monthly plans |
| Free credits | 3,000 once, no card | 1,000 every month, no card |
| Firecrawl figures last verified against their published pricing. Prices change; if one is stale, tell us. | ||
Limits worth knowing before you crawl
These describe how the crawler behaves today, as read from its code. Where a number is a default, it is marked as one.
- 5,000 pages per job, and a 30-minute run time by default. A job that hits the clock stops with stopReason "job deadline reached", so split big sites by path.
- No resume: re-running a crawl starts a new job, and its pages are fetched and billed again.
- Results are kept for 7 days by default, then expire.
- Every page is a live fetch; crawls do not read the scrape cache.
- No logins: the crawler sends no cookies or auth headers.
- robots.txt is obeyed by default and requests to one host are staggered; Crawl-delay itself is not read.
- Images, scripts, archives and feeds are skipped, and tracking parameters such as utm_* and gclid are stripped, so each page is fetched once.
- The webhook is sent once, when the job ends. If your endpoint was down, poll the status URL.
- /map reads a capped number of sitemap files per call; on a site with no sitemap it falls back to the homepage's links, a sample rather than an inventory.
What teams crawl sites for
- Loading a docs site or help centre into a RAG index, one markdown page per record
- Building an LLM knowledge base from a company's own public pages
- Pre-migration audits: every URL with its title, status code, word count and canonical
- Internal link graphs with storeContent: false and formats ["links"]
- Weekly competitor monitoring: re-crawl /blog/** and diff titles and word counts
- Finding thin or broken pages by statusCode and wordCount
When a search beats a crawl
A crawl answers "what is on this one site". For "which pages across the web answer this query", POST /search-and-scrape runs a web search and scrapes the result pages in one call: 1 credit for the search plus 1 per page scraped, so five results with their full text cost 6 credits ($0.0018 at Scale-pack rates).
FAQ
Web crawler API FAQ
What is a web crawler API?
A web crawler API crawls a website for you: it starts from one URL, follows the site's links and sitemaps within your rules, fetches each page through managed infrastructure and returns the pages as data. Searlo's runs as a background job covering up to 5,000 pages, with no crawler, queue or proxy pool for you to operate.
How much does it cost to crawl a website?
1 credit per page actually crawled: $0.30 per 1,000 pages at the Scale pack, from $0.20 (Enterprise) to $0.80 (Micro). Failed pages, robots-disallowed pages and pages beyond your limit cost nothing, and reading status or results is free.
How do I crawl only part of a site?
Pass includePaths and excludePaths, up to 25 globs each, matched against the URL path. * matches within one path segment, ** matches across segments, and a pattern with no wildcard is a prefix, so /docs keeps everything under /docs. An include list is a whitelist: anything outside it is skipped.
Does the crawler respect robots.txt?
Yes, by default. It reads robots.txt before the first page, skips disallowed paths (shown as skipped, not billed) and staggers requests to the same host. Set respectRobots: false only for a site you own or have permission to crawl.
How do I know when a crawl has finished?
Poll GET /crawl/{jobId} until status is completed, failed or cancelled, or pass webhook: { url, secret } for a single POST with the job's progress, creditsCharged and resultsUrl. With a secret, the POST carries x-searlo-signature: v1= plus the hex HMAC-SHA256 of the x-searlo-timestamp value, a dot and the raw body.
Can it crawl JavaScript-rendered sites?
Yes. A page that comes back as an empty client-side shell gets one pass in a real browser automatically, still at 1 credit; it just takes longer. Set render: true to render every page, or render: false to never use the browser.
What is the difference between /map and /crawl?
/map answers "which URLs exist" from the site's own robots.txt and sitemaps, without fetching any page content: 1 credit for up to 5,000 URLs. /crawl fetches the pages themselves and costs 1 credit per page. Run /map first to size a crawl, or to build a URL list for /extract.
Is it a good site crawler API for LLM pipelines?
It returns what LLM pipelines usually need: main-content markdown with navigation and footers stripped, plus title, description, language, canonicalUrl and wordCount for every page. For token-sized chunks, scrape individual pages with chunk: true through the URL to markdown API.
Crawl your first site free
3,000 free credits and no card: enough to map a site and crawl nearly 3,000 of its pages.