URL to Markdown API
URL to Markdown API for LLM and RAG ingestion
POST a URL to /scrape and get the page back as clean markdown: main content only, with navigation, headers, footers and scripts removed, plus optional token-sized chunks for embedding. It costs 1 credit per page, or 2 when you pass maxAge or cache:false and the page has to be fetched live.
{ "searchTime": 0.17, "totalResults": 85000, "page": 1, "organic": [ { "position": 1, "title": "31.10. Connection Pools and Data Sources", "link": "www.postgresql.org/docs/…", "domain": "postgresql.org", "snippet": "For an environment without an application se…" }, { "position": 2, "domain": "medium.com", … }, { "position": 3, "domain": "stackoverflow.blog", … } ], "credits": 1}How do I convert a webpage to markdown with an API?
Send POST https://api.searlo.tech/api/v1/scrape with the x-api-key header and a body of {"url": "…", "formats": ["markdown"]}. The response's markdown field holds the page's main content; add "chunk": true to get token-sized chunks as well. It costs 1 credit, which is $0.30 per 1,000 pages at the Scale pack.
The price is 1 credit whether the page comes from cache or from the site. The second credit applies only when you pass maxAge or cache:false and the answer is fetched live rather than served from a copy young enough for you. Scrapes are cached for 6 hours by default, and a fetch that fails is not billed.
- Price
- 1 credit per URL: $0.30 per 1,000 pages at the Scale pack— Searlo pricing
- Forced live fetch
- 2 credits, when maxAge or cache:false sends the call to the site
- Output
- markdown, text, html, links, images, headings, tables, json_ld: any mix, one price
- Chunks
- chunkMaxTokens from 256 to 16,384; 4,096 by default
- Free to start
- 3,000 credits, no card— Searlo pricing
- Firecrawl basic scrape
- $0.599 to $3.20 per 1,000 pages by plan, billed annually— firecrawl.dev/pricing
Third-party figures on this page last verified against the sources linked above. Prices change; if you find one of these stale, tell us and we will correct it.
Language models read markdown well and raw HTML badly. A typical page carries scripts, styles, navigation, cookie banners and footer links that cost tokens and add nothing to an answer. Converting it to markdown keeps what a model needs (headings, paragraphs, lists, tables, code and links) and drops the rest, and the heading structure gives a chunker natural places to split.
Searlo's /scrape endpoint does that conversion in one call. It fetches the URL through a managed proxy pool, gives the page one real browser pass if it arrives as an empty JavaScript shell, isolates the main content, converts it to markdown, and returns it with the page's metadata. Other views of the same page (plain text, cleaned HTML, links, images, the heading outline, tables, the page's own JSON-LD) come back in the same call at no extra cost.
This page covers single-URL conversion. For a whole site, or for typed fields such as price or author instead of prose, see the related Web Data pages further down.
What the markdown keeps, and what it drops
Main content only
With onlyMainContent (the default) the article or main region is kept. Nav, header, footer, aside, forms, scripts and iframes are removed first.
Structure a model can use
# headings, nested lists, blockquotes, bold and italic, and fenced code blocks, so chunkers can split on sections.
Links and images, your call
includeLinks and includeImages switch inline [text](url) links and  images on or off.
Token-sized chunks
chunk: true splits on paragraph boundaries with a small overlap. Set chunkMaxTokens between 256 and 16,384.
Metadata for citations
title, description, author, publishedDate, language, canonicalUrl, wordCount and Open Graph tags come back with every page.
JavaScript pages handled
An empty client-side shell gets one browser pass automatically. The via field says http or browser.
Convert a URL to markdown
One POST per page. The second cURL call shows the freshness dial: maxAge is the oldest cached copy you will accept, in seconds.
# Convert one page to markdown, split into ~1,024-token chunks (1 credit)
curl -X POST "https://api.searlo.tech/api/v1/scrape" \
-H "x-api-key: YOUR_API_KEY" \
-H "content-type: application/json" \
-d '{
"url": "https://example.com/blog/how-we-ship",
"formats": ["markdown"],
"onlyMainContent": true,
"excludeSelectors": [".newsletter-signup", "#comments"],
"chunk": true,
"chunkMaxTokens": 1024
}'
# Accept a cached copy only if it is under an hour old.
# Served from cache: 1 credit. Fetched live because of maxAge: 2 credits.
curl -X POST "https://api.searlo.tech/api/v1/scrape" \
-H "x-api-key: YOUR_API_KEY" \
-H "content-type: application/json" \
-d '{ "url": "https://example.com/pricing", "formats": ["markdown"], "maxAge": 3600 }'import requests
r = requests.post(
"https://api.searlo.tech/api/v1/scrape",
headers={"x-api-key": "YOUR_API_KEY"},
json={
"url": "https://example.com/blog/how-we-ship",
"formats": ["markdown"],
"chunk": True,
"chunkMaxTokens": 1024,
},
timeout=90,
)
r.raise_for_status()
page = r.json()
meta = page["metadata"]
source = meta["canonicalUrl"] or page["url"]
print(meta["title"], "|", meta["wordCount"], "words |", page["via"])
print("cache:", r.headers.get("X-Cache"), "| credits:", page["credits"])
# One record per chunk, each carrying its citation back to the page.
records = [
{
"id": f"{source}#{chunk['index']}",
"text": chunk["text"],
"tokens_estimate": chunk["tokens"],
"source": source,
"title": meta["title"],
}
for chunk in page.get("chunks", [])
]
print(len(records), "chunks ready to embed")const res = await fetch("https://api.searlo.tech/api/v1/scrape", {
method: "POST",
headers: { "x-api-key": "YOUR_API_KEY", "content-type": "application/json" },
body: JSON.stringify({
url: "https://example.com/blog/how-we-ship",
formats: ["markdown"],
chunk: true,
chunkMaxTokens: 1024,
}),
});
if (!res.ok) throw new Error(`scrape failed: ${res.status}`); // failures are not billed
const page = await res.json();
console.log(page.metadata.title, page.via, res.headers.get("x-cache"), page.credits);
for (const chunk of page.chunks ?? []) {
// tokens is an estimate (about 1.3 per word), not your model's tokenizer.
console.log(`#${chunk.index} ~${chunk.tokens} tokens`, chunk.text.slice(0, 80));
}Response (illustrative)
Field names are the API's own; values are illustrative and the markdown is trimmed. links comes back unless you set includeLinks: false. The X-Cache (HIT or MISS) and X-Credits-Deducted response headers say where the page came from and what the call cost.
{
"searchParameters": {
"url": "https://example.com/blog/how-we-ship",
"type": "scrape",
"formats": ["markdown"]
},
"url": "https://example.com/blog/how-we-ship",
"requestedUrl": "https://example.com/blog/how-we-ship",
"statusCode": 200,
"via": "http",
"bytes": 148213,
"markdown": "# How we ship\n\nWe deploy small changes many times a day...\n\n## Every change starts as a draft\n\n...",
"chunks": [
{ "index": 0, "text": "# How we ship\n\nWe deploy small changes many times a day...", "tokens": 1012 },
{ "index": 1, "text": "## Review happens in the open\n\n...", "tokens": 987 }
],
"metadata": {
"title": "How we ship | Example Blog",
"description": "The release process behind a small engineering team.",
"author": "Jane Doe",
"publishedDate": "2026-08-14T09:00:00Z",
"language": "en",
"canonicalUrl": "https://example.com/blog/how-we-ship",
"robots": null,
"wordCount": 1874,
"openGraph": { "title": "How we ship", "type": "article", "image": "https://example.com/og/how-we-ship.png" },
"twitter": null
},
"links": [
{ "url": "https://example.com/blog", "text": "Blog", "rel": null, "internal": true }
],
"source": "searlo.tech(scrape)",
"credits": 1
}Request parameters
| Parameter | Default | What it does |
|---|---|---|
| url | required | The page to convert: http(s), up to 2,048 characters |
| formats | markdown | Any of markdown, text, html, links, images, headings, tables, json_ld, in one call |
| onlyMainContent | true | Keep the main content region; false converts the whole body |
| includeLinks / includeImages | true | Inline links and images in the markdown |
| excludeSelectors | none | Up to 20 CSS selectors to remove before conversion |
| chunk / chunkMaxTokens | false / 4,096 | Split the markdown into chunks of 256 to 16,384 estimated tokens |
| render | automatic | true forces a browser render; false never uses one |
| waitForSelector / timeout | none | Wait for an element when the page is rendered in a browser; timeout from 1,000 to 60,000 ms |
| actions | none | Up to 20 in-page steps (click, type, press, scroll, wait, waitForSelector, evaluate) before reading; forces a browser |
| maxAge | none | Oldest cached copy you accept, in seconds (0 = always live); a live fetch it forces costs 2 credits |
| cache | true | false is the same as maxAge: 0 |
What conversion costs
| Pack | Price | Pages at 1 credit | Pages at 2 credits | Per 1,000 pages |
|---|---|---|---|---|
| Micro | $3.99 | 5,000 | 2,500 | $0.80 |
| Starter | $9.99 | 20,000 | 10,000 | $0.50 |
| Builder | $29.99 | 75,000 | 37,500 | $0.40 |
| Scale | $74.99 | 250,000 | 125,000 | $0.30 |
| Pro | $199.99 | 900,000 | 450,000 | $0.22 |
| Enterprise | $799 | 4,000,000 | 2,000,000 | $0.20 |
Pricing maths
Converting a 2,000-page knowledge base once costs 2,000 credits: $0.60 at Scale-pack rates, and it fits inside the 3,000 free credits.
Re-converting the same 2,000 pages every day for 30 days is 60,000 credits, $18.00 at Scale-pack rates, as long as you leave maxAge and cache unset. Add maxAge: 3600 to that job and every page the cache cannot answer within the hour costs 2 credits. Cached copies last 6 hours by default, so a once-a-day refresh with maxAge set would be all live fetches at double the price: $36.00.
The rule of thumb: pass maxAge or cache:false only when you need a copy younger than the cache would otherwise give you. A plain cache miss is never surcharged.
Need search results with their page content? POST /search-and-scrape runs a web search and scrapes the result pages in one call, for 1 credit for the search plus 1 per page scraped. Five results with their markdown cost 6 credits, or $0.0018 at Scale-pack rates.
Searlo vs Firecrawl for URL to markdown
Both price a basic markdown scrape at one credit per page; the difference is what a credit costs. Firecrawl's figures are from firecrawl.dev/pricing, checked 2026-09-29. Firecrawl and Google are trademarks of their owners; Searlo is independent and affiliated with neither.
| Capability | Searlo | Firecrawl |
|---|---|---|
| Price per 1,000 pages (basic scrape) | $0.30 at the Scale pack; $0.20–$0.80 by pack | $0.599 (Scale) to $3.20 (Hobby), billed annually |
| Free credits | 3,000 once, no card | 1,000 every month, no card |
| Billing model | One-time credit packs | Monthly plans |
| Structured JSON from the page | /extract: 1 credit per successful URL | The JSON format adds 4 credits per page |
| Firecrawl figures last verified against their published pricing. Prices change; if one is stale, tell us. | ||
Where HTML-to-markdown conversion falls short
The converter is built for clean LLM input, not for pixel-perfect fidelity. Know these limits before you depend on it:
- Main-content detection is a heuristic: the first article, main or common content container with real text, else the densest text block. On unusual layouts, set onlyMainContent: false and trim noise with excludeSelectors.
- Tables come out as pipe-separated rows without a header divider, which strict Markdown parsers may not read as tables. The tables format gives headers and rows as arrays.
- Fenced code blocks carry no language tag.
- Inline links keep the page's own href, which may be relative. The links format returns absolute URLs.
- Chunk sizes are an estimate (about 1.3 tokens per word), not your model's tokenizer. A paragraph longer than chunkMaxTokens stays whole, and the overlap is fixed at about 128 tokens.
- HTML pages only: PDFs and other files are not converted.
- No logins: /scrape sends no cookies or auth headers.
- A page that answers 404 or shows an error message is still a completed fetch and costs a credit, so check statusCode.
- A browser render adds seconds. Set timeout (up to 60,000 ms) and give your HTTP client a longer one.
Where teams use URL to markdown
- RAG ingestion: markdown chunks with the canonical URL attached for citations
- An agent's "read this page" tool, one call per link the agent decides to open
- Summarising articles without the site chrome around them
- Change monitoring: snapshot pages as markdown and diff them
- Moving an HTML knowledge base into a markdown docs system
- Search, then read: /search-and-scrape for query-driven research
Related Web Data pages
Web scraping API overview
Scrape, map, crawl and extract on one key, with the full pricing comparison.
Web crawler API
Turn a whole site into markdown pages as one background job.
Data extraction API
Typed fields instead of prose, read from the page's own structured data.
Search API for RAG
Live search results to decide which pages are worth converting.
SERP API for LLMs
Ground answers in Google results, then read the pages behind them.
Firecrawl alternative
Scrape pricing and features side by side.
FAQ
URL to markdown API FAQ
What is a URL to markdown API?
An API that takes a web address and returns that page's readable content as markdown instead of raw HTML. Searlo's fetches the page (with a browser pass when the page needs JavaScript), strips navigation, scripts and other page furniture, and returns markdown plus metadata such as the title, author and canonical URL.
How much does it cost to convert a webpage to markdown?
1 credit per URL: $0.30 per 1,000 pages at the Scale pack ($74.99), from $0.20 on the Enterprise pack to $0.80 on the Micro pack. Extra formats in the same call, such as text, html, links or tables, are included. A call costs 2 credits only when you pass maxAge or cache:false and the page is fetched live.
Why does a scrape sometimes cost 2 credits?
Because you asked for freshness and the cache could not provide it. With maxAge or cache:false set, an answer served from a cached copy young enough for you stays at 1 credit, and one fetched live from the site costs 2. Without those parameters every call is 1 credit, cached or not. The credits field and the X-Credits-Deducted header show what each call cost.
Does it work on JavaScript-heavy pages?
Yes. If the plain fetch returns an empty client-side shell, the page gets one pass in a real browser automatically, at the same price. Use render: true to always render (with waitForSelector to wait for a specific element), or render: false to never use the browser.
How should I chunk the markdown for embeddings?
Set chunk: true and a chunkMaxTokens between 256 and 16,384. The markdown is split on paragraph boundaries with about 128 tokens of overlap, and each chunk comes back with its index and an estimated token count. The estimate is about 1.3 tokens per word, so leave headroom below your embedding model's limit.
Can I convert an entire website to markdown?
Yes, with POST /crawl instead of /scrape. It walks the site as a background job and stores each page's markdown, at 1 credit per page crawled. The web crawler API page covers depth, path filters and webhooks.
Is markdown cheaper for an LLM than HTML?
Usually, because scripts, styles, attributes and navigation are removed before conversion, and those make up much of a page's HTML. How much you save depends on the page. The wordCount field and each chunk's token estimate let you measure it on your own pages instead of trusting a headline figure.
Is there a GET version for agents?
Yes. GET /search/webpage?url=… returns one page's readable content (markdown by default) in a content field, with its title, description, author and published date, also for 1 credit. /scrape is the richer call: several formats at once, chunks, selectors and the freshness dial.
Convert your first 3,000 pages free
3,000 free credits on signup, no card. At 1 credit a page, that is 3,000 pages of markdown.