Buyer's guide
The best HTML-to-Markdown APIs for LLMs & RAG
Feeding web pages to a language model works best when the HTML becomes clean Markdown first. These 8 APIs and tools do it in 2026 — what each is best for, what it costs, how much junk its approach lets through, and the honest trade-offs.
By Roger Campos · Last updated: September 2026
TL;DR
For per-URL Markdown with screenshots, metadata or audits from the same key, URLpipe — its deterministic converter left 11.0% hidden text on our 33-page benchmark, against 17.0–21.3% for the strategies competitors use. For crawling whole sites, Firecrawl or Apify. For the fastest one-off call or PDFs, Jina Reader. Microlink and Urlbox add Markdown to link previews and screenshots; ScrapingBee brings proxies; Crawl4AI is the open-source, self-hosted option.
Free plan, no credit card. 1,000 credits a month.
What to look for
How to choose an HTML-to-Markdown API
Before comparing tools, weigh these six criteria against your own use case. The right pick usually comes down to scope (one URL or a whole site) and how much data you need beyond text.
Markdown quality, measured
Does it keep the article and drop the chrome? Ask for two numbers: how much visible text survives, and how much of the output is text a visitor never saw — menus, hidden panels, cookie policies.
JavaScript rendering
Modern sites are client-rendered. Without a real browser you get an empty shell for a single-page app.
Scope
One URL at a time, or a full-site crawl? Per-URL APIs are simpler; crawlers are built for breadth.
Data beyond Markdown
You may also want metadata, a screenshot, a summary or an audit of the same page. One vendor is one integration.
What a call costs
Credits per page, tokens per response, or renders? Check what a cache hit and a failed fetch cost, and what happens at the quota — a 402, a pause, or overage.
Agents, hosting and licensing
A hosted MCP server, a data region, a self-hosting licence — any of these can be a hard requirement.
Measured
How much junk each approach lets through
We ran four main-content strategies over the same 33 real pages, each on the same fetched DOM, and measured two things: how much of the text a visitor sees survives, and how much of the output is text a visitor never sees.
| Strategy | Visible text kept | Text the reader never sees | Cost / page | Time / page |
|---|---|---|---|---|
| URLpipe (deterministic DOM walk) | 87.2% | 11.0% | $0 | ~20 ms |
| Readability (the approach Jina Reader runs) | 88.2% | 21.3% | — | — |
| Firecrawl's selector blocklist | 85.8% | 17.0% | — | — |
| LLM conversion (a model rewrites the page) | 79.0% | 22.0% | $0.0025 | 39 s |
These are the strategies reimplemented on one DOM, not the vendors' live APIs, which add their own tuning. Coverage is close across the board; the difference is the second column, which is what ends up in your chunks. URLpipe's converter reads no CSS, so text hidden only by a stylesheet can still leak.
At a glance
The 8 best HTML-to-Markdown APIs compared
| Tool | Renders JS | Full-site crawl | Free tier | Standout |
|---|---|---|---|---|
| URLpipe | Yes | No | 1,000 credits/mo | Measured output, 8 data types |
| Firecrawl | Yes | Yes | 1,000 credits/mo | Crawl + schema extraction |
| Jina Reader | Yes | No | 20 RPM + 10M tokens | Zero-signup prefix |
| Microlink | Yes | No | 25 requests/day | Markdown + link previews |
| Urlbox | Yes | No | 7-day trial | Screenshot + Markdown in one render |
| ScrapingBee | Yes | No | 1,000 trial credits | Proxies + extraction rules |
| Apify | Yes | Yes | $5 credit/mo | Actor marketplace |
| Crawl4AI | Yes | Yes | Free (self-host) | Self-hosted, open source |
“Full-site crawl” marked partial means single-URL first with limited multi-page support. Pricing and limits change — confirm current numbers on each vendor's site.
The tools
Each option in depth
URLpipe first — it's ours, so read its watch-out as carefully as the others'. Every entry has what it's best for, pricing checked on the vendor's own site on September 24, 2026, and an honest watch-out.
URLpipe
That's usOne API, eight kinds of clean data from any URL — Markdown made without a model.
- Best for
- Per-URL Markdown for models and agents, when you also want screenshots, metadata or audits.
- Pricing
- Free: 1,000 credits a month, no card. Paid from $19/mo, never cut off.
Strengths
- Markdown from a deterministic walk of the rendered DOM: on our 33-page benchmark, 11.0% of the output was text a visitor never sees, against 17.0–21.3% for the strategies competitors use
- 1 credit per page whatever its length; cache hits free for up to 30 days
- Screenshots, metadata, summaries, keywords, Lighthouse and console errors from the same key — /scrape takes several from one page visit
- Hosted MCP server, so an agent can read JavaScript-rendered pages with nothing installed
- Official clients for Python, JavaScript, Ruby and Go, plus LangChain and LlamaIndex packages and an n8n node
- Async by default with signed webhooks; idempotency keys and labels
- Fetched, rendered and stored in the EU; AI processing EU-only at the flip of a setting
Watch out
Single URLs only — no crawling, no URL discovery, no schema extraction, no PDFs. The converter reads no CSS, so text hidden only by a stylesheet can slip through. Credits are weighted: Markdown is 1, an AI summary 17.
Firecrawl
The crawl-and-extract API for taking in whole sites.
- Best for
- Crawling whole sites, schema-based extraction and PDFs.
- Pricing
- Free: 1,000 credits/mo. Hobby $16/mo annual (5,000 credits) up to Scale $599/mo annual (1M).
Strengths
- Crawl, map, batch scrape and schema-based JSON extraction
- Several formats from one scrape — Markdown, HTML, links, screenshot, summary
- PDF and document parsing
- Official SDKs in nine languages and a hosted MCP server
- AGPL-3.0 core you can self-host
Watch out
About one credit per page, +4 for JSON extraction, and cached results still spend credits. At quota the API answers 402 unless auto-recharge is on. Its main-content strategy is a selector blocklist: on our benchmark that approach left 17.0% text a visitor never sees.
Jina Reader
Prepend r.jina.ai/ to any URL and get Markdown back.
- Best for
- The lowest-friction Markdown call, PDFs and image captions.
- Pricing
- Free: 20 req/min with no key; 10M free tokens per new key. Then billed by output tokens.
Strengths
- No-signup URL prefix — try any URL in a browser
- PDF reading and image captions from a vision model
- CSS selectors to target or remove elements, JSON-schema extraction
- Hosted MCP server; an experimental EU endpoint
- Apache-2.0 repository you can self-host
Watch out
Token billing grows with page length and the cache lasts about five minutes. Owned by Elastic since October 2025; the ReaderLM-v2 model is CC-BY-NC. On our benchmark the Readability approach it runs left 21.3% text a visitor never sees.
Microlink
A metadata and browser API that returns Markdown on every plan.
- Best for
- Link previews and screenshots, with Markdown from the same request.
- Pricing
- Free: 25 requests/day, no card. Pro $49/mo for 46,000 requests.
Strengths
- Markdown, metadata, screenshots, PDF and a Lighthouse report as flags on one request
- Markdown scoped to a selector, and from PDF and Office documents
- Cache hits don't count against the quota
- Pauses at the quota instead of billing overage
- Open-source metascraper (MIT) behind the metadata
Watch out
The free tier is 25 requests a day, and paid plans pause at the ceiling until the next cycle. No AI summary, keyword or async webhook option. Its MCP server runs locally and its repository was archived in July 2026.
Urlbox
A rendering API — screenshots, PDFs and video — that also saves Markdown.
- Best for
- Teams already rendering screenshots who want Markdown off the same render.
- Pricing
- No free plan; 7-day trial. Markdown from Hi-Fi, $49/mo for 5,000 renders.
Strengths
- One render returns an image plus HTML, Markdown and metadata
- Async renders with signed webhooks
- Deep screenshot options, PDF and video
- LLM extraction against a schema with your own model key (Ultra and up)
Watch out
Markdown and metadata start at Hi-Fi, not the entry Lo-Fi plan. Priced per render, so Markdown costs what a screenshot costs. No free tier, and no proxies of its own.
ScrapingBee
A scraping API built around proxies and getting the page at all.
- Best for
- Sites that need premium proxies and geotargeting.
- Pricing
- 1,000 free trial credits. Hobby $19/mo (75,000 credits) to Business+ $599/mo.
Strengths
- Premium (residential) and stealth proxies, country targeting
- Page Markdown or plain text, screenshots, CSS and AI extraction rules
- Only successful requests are billed
- Hosted MCP server
Watch out
A JavaScript-rendered request costs 5 credits, 25 with a premium proxy and 75 in stealth mode, so plan sizes shrink fast. Synchronous only, 140-second timeout, and no automatic overage.
Apify
A platform of scrapers (Actors); Website Content Crawler outputs Markdown for RAG.
- Best for
- Crawling documentation or whole sites into a vector store, on a schedule.
- Pricing
- Free: $5 of platform credit/mo. Starter $19/mo. Actors bill by compute.
Strengths
- Website Content Crawler: Markdown, HTML, metadata and linked documents across a site
- Scheduling, storage, retries and webhooks on the platform
- Integrations with LangChain, LlamaIndex, Pinecone and Qdrant
- Hosted MCP server at mcp.apify.com
Watch out
You pay for compute rather than per page — Apify estimates $0.50–$5 per 1,000 pages with a headless browser — so costs are harder to predict. A platform to learn, not one endpoint.
Crawl4AI
Open-source Python crawler with LLM-ready Markdown.
- Best for
- Teams that want to self-host with no per-page fee.
- Pricing
- Free and open source (Apache-2.0). You pay for the servers it runs on.
Strengths
- Apache-2.0 — commercial use without a licence fee
- Clean and “fit” Markdown, screenshots, PDFs, CSS/XPath or LLM extraction
- Deep crawling with resume; a Docker server with MCP support
- Full control of the browser stack
Watch out
You run, scale and unblock it yourself. The hosted Crawl4AI Cloud API is a closed beta.
Recommendations
Which one should you pick?
Feed single pages to a model or agent, with screenshots, metadata or audits alongside
URLpipeCrawl and structure an entire website for a RAG index
Firecrawl or ApifyQuickest possible one-off Markdown, or reading PDFs
Jina ReaderLink previews at volume, with Markdown on the side
MicrolinkAlready rendering screenshots and want Markdown from the same render
UrlboxSites that need premium proxies or geotargeting
ScrapingBeeSelf-host with no per-page cost and full control
Crawl4AIFAQ
Frequently asked questions
What is the best HTML-to-Markdown API?
How should I choose a URL-to-Markdown API for RAG?
Why convert HTML to Markdown for LLMs?
Does the conversion use an AI model?
Which HTML-to-Markdown API has the best free tier?
Is there a good open-source option?
Try the measured one.
URLpipe turns any URL into Markdown, screenshots, metadata, summaries, keywords, Lighthouse audits and console errors — over HTTP or MCP. 1,000 credits a month, no card.