Skip to main content

Buyer's guide

The best HTML-to-Markdown APIs for LLMs & RAG

Feeding web pages to a language model works best when the HTML becomes clean Markdown first. These 8 APIs and tools do it in 2026 — what each is best for, what it costs, how much junk its approach lets through, and the honest trade-offs.

By · Last updated: September 2026

TL;DR

For per-URL Markdown with screenshots, metadata or audits from the same key, URLpipe — its deterministic converter left 11.0% hidden text on our 33-page benchmark, against 17.0–21.3% for the strategies competitors use. For crawling whole sites, Firecrawl or Apify. For the fastest one-off call or PDFs, Jina Reader. Microlink and Urlbox add Markdown to link previews and screenshots; ScrapingBee brings proxies; Crawl4AI is the open-source, self-hosted option.

Free plan, no credit card. 1,000 credits a month.

What to look for

How to choose an HTML-to-Markdown API

Before comparing tools, weigh these six criteria against your own use case. The right pick usually comes down to scope (one URL or a whole site) and how much data you need beyond text.

Markdown quality, measured

Does it keep the article and drop the chrome? Ask for two numbers: how much visible text survives, and how much of the output is text a visitor never saw — menus, hidden panels, cookie policies.

JavaScript rendering

Modern sites are client-rendered. Without a real browser you get an empty shell for a single-page app.

Scope

One URL at a time, or a full-site crawl? Per-URL APIs are simpler; crawlers are built for breadth.

Data beyond Markdown

You may also want metadata, a screenshot, a summary or an audit of the same page. One vendor is one integration.

What a call costs

Credits per page, tokens per response, or renders? Check what a cache hit and a failed fetch cost, and what happens at the quota — a 402, a pause, or overage.

Agents, hosting and licensing

A hosted MCP server, a data region, a self-hosting licence — any of these can be a hard requirement.

Measured

How much junk each approach lets through

We ran four main-content strategies over the same 33 real pages, each on the same fetched DOM, and measured two things: how much of the text a visitor sees survives, and how much of the output is text a visitor never sees.

StrategyVisible text keptText the reader never seesCost / pageTime / page
URLpipe (deterministic DOM walk)87.2%11.0%$0~20 ms
Readability (the approach Jina Reader runs)88.2%21.3%——
Firecrawl's selector blocklist85.8%17.0%——
LLM conversion (a model rewrites the page)79.0%22.0%$0.002539 s

These are the strategies reimplemented on one DOM, not the vendors' live APIs, which add their own tuning. Coverage is close across the board; the difference is the second column, which is what ends up in your chunks. URLpipe's converter reads no CSS, so text hidden only by a stylesheet can still leak.

At a glance

The 8 best HTML-to-Markdown APIs compared

ToolRenders JSFull-site crawlFree tierStandout
URLpipeYesNo1,000 credits/moMeasured output, 8 data types
FirecrawlYesYes1,000 credits/moCrawl + schema extraction
Jina ReaderYesNo20 RPM + 10M tokensZero-signup prefix
MicrolinkYesNo25 requests/dayMarkdown + link previews
UrlboxYesNo7-day trialScreenshot + Markdown in one render
ScrapingBeeYesNo1,000 trial creditsProxies + extraction rules
ApifyYesYes$5 credit/moActor marketplace
Crawl4AIYesYesFree (self-host)Self-hosted, open source

“Full-site crawl” marked partial means single-URL first with limited multi-page support. Pricing and limits change — confirm current numbers on each vendor's site.

The tools

Each option in depth

URLpipe first — it's ours, so read its watch-out as carefully as the others'. Every entry has what it's best for, pricing checked on the vendor's own site on September 24, 2026, and an honest watch-out.

1

URLpipe

That's us

One API, eight kinds of clean data from any URL — Markdown made without a model.

Best for
Per-URL Markdown for models and agents, when you also want screenshots, metadata or audits.
Pricing
Free: 1,000 credits a month, no card. Paid from $19/mo, never cut off.

Strengths

  • Markdown from a deterministic walk of the rendered DOM: on our 33-page benchmark, 11.0% of the output was text a visitor never sees, against 17.0–21.3% for the strategies competitors use
  • 1 credit per page whatever its length; cache hits free for up to 30 days
  • Screenshots, metadata, summaries, keywords, Lighthouse and console errors from the same key — /scrape takes several from one page visit
  • Hosted MCP server, so an agent can read JavaScript-rendered pages with nothing installed
  • Official clients for Python, JavaScript, Ruby and Go, plus LangChain and LlamaIndex packages and an n8n node
  • Async by default with signed webhooks; idempotency keys and labels
  • Fetched, rendered and stored in the EU; AI processing EU-only at the flip of a setting

Watch out

Single URLs only — no crawling, no URL discovery, no schema extraction, no PDFs. The converter reads no CSS, so text hidden only by a stylesheet can slip through. Credits are weighted: Markdown is 1, an AI summary 17.

2

Firecrawl

The crawl-and-extract API for taking in whole sites.

Best for
Crawling whole sites, schema-based extraction and PDFs.
Pricing
Free: 1,000 credits/mo. Hobby $16/mo annual (5,000 credits) up to Scale $599/mo annual (1M).

Strengths

  • Crawl, map, batch scrape and schema-based JSON extraction
  • Several formats from one scrape — Markdown, HTML, links, screenshot, summary
  • PDF and document parsing
  • Official SDKs in nine languages and a hosted MCP server
  • AGPL-3.0 core you can self-host

Watch out

About one credit per page, +4 for JSON extraction, and cached results still spend credits. At quota the API answers 402 unless auto-recharge is on. Its main-content strategy is a selector blocklist: on our benchmark that approach left 17.0% text a visitor never sees.

3

Jina Reader

Prepend r.jina.ai/ to any URL and get Markdown back.

Best for
The lowest-friction Markdown call, PDFs and image captions.
Pricing
Free: 20 req/min with no key; 10M free tokens per new key. Then billed by output tokens.

Strengths

  • No-signup URL prefix — try any URL in a browser
  • PDF reading and image captions from a vision model
  • CSS selectors to target or remove elements, JSON-schema extraction
  • Hosted MCP server; an experimental EU endpoint
  • Apache-2.0 repository you can self-host

Watch out

Token billing grows with page length and the cache lasts about five minutes. Owned by Elastic since October 2025; the ReaderLM-v2 model is CC-BY-NC. On our benchmark the Readability approach it runs left 21.3% text a visitor never sees.

4

Microlink

A metadata and browser API that returns Markdown on every plan.

Best for
Link previews and screenshots, with Markdown from the same request.
Pricing
Free: 25 requests/day, no card. Pro $49/mo for 46,000 requests.

Strengths

  • Markdown, metadata, screenshots, PDF and a Lighthouse report as flags on one request
  • Markdown scoped to a selector, and from PDF and Office documents
  • Cache hits don't count against the quota
  • Pauses at the quota instead of billing overage
  • Open-source metascraper (MIT) behind the metadata

Watch out

The free tier is 25 requests a day, and paid plans pause at the ceiling until the next cycle. No AI summary, keyword or async webhook option. Its MCP server runs locally and its repository was archived in July 2026.

5

Urlbox

A rendering API — screenshots, PDFs and video — that also saves Markdown.

Best for
Teams already rendering screenshots who want Markdown off the same render.
Pricing
No free plan; 7-day trial. Markdown from Hi-Fi, $49/mo for 5,000 renders.

Strengths

  • One render returns an image plus HTML, Markdown and metadata
  • Async renders with signed webhooks
  • Deep screenshot options, PDF and video
  • LLM extraction against a schema with your own model key (Ultra and up)

Watch out

Markdown and metadata start at Hi-Fi, not the entry Lo-Fi plan. Priced per render, so Markdown costs what a screenshot costs. No free tier, and no proxies of its own.

6

ScrapingBee

A scraping API built around proxies and getting the page at all.

Best for
Sites that need premium proxies and geotargeting.
Pricing
1,000 free trial credits. Hobby $19/mo (75,000 credits) to Business+ $599/mo.

Strengths

  • Premium (residential) and stealth proxies, country targeting
  • Page Markdown or plain text, screenshots, CSS and AI extraction rules
  • Only successful requests are billed
  • Hosted MCP server

Watch out

A JavaScript-rendered request costs 5 credits, 25 with a premium proxy and 75 in stealth mode, so plan sizes shrink fast. Synchronous only, 140-second timeout, and no automatic overage.

7

Apify

A platform of scrapers (Actors); Website Content Crawler outputs Markdown for RAG.

Best for
Crawling documentation or whole sites into a vector store, on a schedule.
Pricing
Free: $5 of platform credit/mo. Starter $19/mo. Actors bill by compute.

Strengths

  • Website Content Crawler: Markdown, HTML, metadata and linked documents across a site
  • Scheduling, storage, retries and webhooks on the platform
  • Integrations with LangChain, LlamaIndex, Pinecone and Qdrant
  • Hosted MCP server at mcp.apify.com

Watch out

You pay for compute rather than per page — Apify estimates $0.50–$5 per 1,000 pages with a headless browser — so costs are harder to predict. A platform to learn, not one endpoint.

8

Crawl4AI

Open-source Python crawler with LLM-ready Markdown.

Best for
Teams that want to self-host with no per-page fee.
Pricing
Free and open source (Apache-2.0). You pay for the servers it runs on.

Strengths

  • Apache-2.0 — commercial use without a licence fee
  • Clean and “fit” Markdown, screenshots, PDFs, CSS/XPath or LLM extraction
  • Deep crawling with resume; a Docker server with MCP support
  • Full control of the browser stack

Watch out

You run, scale and unblock it yourself. The hosted Crawl4AI Cloud API is a closed beta.

Recommendations

Which one should you pick?

Feed single pages to a model or agent, with screenshots, metadata or audits alongside

URLpipe

Crawl and structure an entire website for a RAG index

Firecrawl or Apify

Quickest possible one-off Markdown, or reading PDFs

Jina Reader

Link previews at volume, with Markdown on the side

Microlink

Already rendering screenshots and want Markdown from the same render

Urlbox

Sites that need premium proxies or geotargeting

ScrapingBee

Self-host with no per-page cost and full control

Crawl4AI

FAQ

Frequently asked questions

What is the best HTML-to-Markdown API?
It depends on scope. For turning individual URLs into Markdown — with screenshots, metadata, summaries or Lighthouse from the same key — URLpipe is the most complete per-URL API, and its Markdown leaks the least hidden text of the strategies we measured. For crawling whole sites and schema extraction, Firecrawl. For the lowest-friction call, Jina Reader. For self-hosting, Crawl4AI.
How should I choose a URL-to-Markdown API for RAG?
Test it on your own pages and check two numbers: how much of the visible text survives, and how much of the output is text a reader never sees — navigation, hidden panels, cookie text — because that is what pollutes your chunks. Then check it renders JavaScript, what a cache hit and a failed fetch cost, and whether you need one URL at a time or a crawl.
Why convert HTML to Markdown for LLMs?
Raw HTML is mostly markup, scripts, styles and navigation that waste tokens and distract a model. Markdown keeps the content and its structure — headings, lists, links, tables — in a compact form that prompts well and chunks cleanly for a vector store.
Does the conversion use an AI model?
It varies. URLpipe's /markdown doesn't: it walks the rendered DOM, which costs nothing per page, takes about 20 ms and gives the same output every time. On our benchmark an LLM conversion scored lower on both coverage (79.0%) and hidden text (22.0%), at $0.0025 and 39 seconds a page. Jina also offers ReaderLM-v2, a model for the job, under a non-commercial licence.
Which HTML-to-Markdown API has the best free tier?
The units differ. Jina gives 10 million tokens per new key and a no-signup prefix at 20 requests a minute. Firecrawl and URLpipe give 1,000 credits a month — about 1,000 pages each — but URLpipe's cache hits are free and Firecrawl's are not. Microlink gives 25 requests a day. Crawl4AI is free to self-host.
Is there a good open-source option?
Yes. Crawl4AI is Apache-2.0 and self-hosts on Docker; Jina's Reader repository is Apache-2.0 too; Firecrawl's core is AGPL-3.0. The trade-off is that you run, scale and unblock them yourself.

Try the measured one.

URLpipe turns any URL into Markdown, screenshots, metadata, summaries, keywords, Lighthouse audits and console errors — over HTTP or MCP. 1,000 credits a month, no card.