Introducing Highlights: the context that matters
Backed byY CombinatorCombinator

Ground RAG in {fresh content}

Collect web pages and documents, keep a source with every chunk, and refresh the content that changes. Build your retrieval pipeline with Scrape, Crawl, Parse, and Monitors.

Get API KeyNo card required
Book a Call
Daydream logo
Kovai logo
Passionfroot logo
Orange logo
SendX logo
Klarna logo
Super.com logo
Daydream logo
Kovai logo
Passionfroot logo
Orange logo
SendX logo
Klarna logo
Super.com logo
Daydream logo
Kovai logo
Passionfroot logo
Orange logo
SendX logo
Klarna logo
Super.com logo
Daydream logo
Kovai logo
Passionfroot logo
Orange logo
SendX logo
Klarna logo
Super.com logo

Collect, index, and keep your content current

Collect web content

Scrape a page or crawl a site. Use Batches when you need to collect a large URL list in the background.

const page = await
  client.web.scrape({
    url: 'https://example.com',
    formats: { markdown: true },
  });

const text = page.markdown.data;

Include your documents

Parse uploaded files into Markdown, then chunk and index them alongside web content with source references.

const document = await
  client.parse.handle(
    createReadStream('report.pdf'),
    { extension: 'pdf' },
  );

const text = document.markdown;

Refresh changed sources

Use Monitor change events to schedule a new scrape and replace the affected chunks in your retrieval index.

const monitor = await
  client.monitors.create({
    name: 'Knowledge base source',
    target: {
      type: 'page',
      url: 'https://example.com',
    },
    change_detection: { type: 'exact' },
    schedule: {
      type: 'interval',
      frequency: 6,
      unit: 'hours',
    },
    webhook: {
      url: 'https://app.example.com/hooks',
      events: ['change.detected'],
    },
  });

Every part of your content pipeline.

Scrape

Collect Markdown from individual pages, keeping each URL with its content for citations and future updates.

POST /v1/web/scrape

Crawl Website

Read a website within your page, depth, and URL filters. Use Batches for larger collections that should run in the background.

POST /v1/web/crawl

Parse File

Convert uploaded documents into Markdown, with optional OCR for scanned PDF pages. Add them to the same index as web content.

POST /v1/parse

Monitor the web

Watch source pages on a schedule. Use signed change events to trigger a new scrape and update affected chunks.

POST /v1/monitors

Use Map to discover a site’s URLs before collecting content, or submit large URL lists to Batches.

Built for

Teams that need live web data in their AI stack

Ship an agent that actually knows things.

Free tier, 10-minute integration, and the same API powering agents at Mintlify, daily.dev, and Propane. No credit card to start.