LlamaIndex

Power up LLM with web scraping

LlamaIndex + Scrapfly. Web data, no code.

  • Official integration. Native LlamaIndex app maintained by Scrapfly, always current with the latest API features.
  • Every Scrapfly product. Web Scraping, Extraction, Screenshot, and MCP, one connection, all capabilities.
  • No code required. Drop Scrapfly actions into your LlamaIndex workflow; connect to 7,000+ other apps.
1,000 free Scrapfly credits. No credit card required.
LlamaIndex Integration
// AT A GLANCE
OAuthone-time auth
4actions
0code required

COVERAGE

LlamaIndex + Every Scrapfly Product

One integration. Four product surfaces. Same engine powering 5B+ monthly scrapes, exposed as LlamaIndex actions.

Every Scrapfly API. Inside LlamaIndex.

One authentication, four product surfaces. Scrape any page, extract structured data, capture screenshots, or drive an agent, all from your LlamaIndex workflow. No SDK install, no secrets rotation plumbing.

4scrapfly APIs exposed
unblocker=Trueone flag bypasses 8 anti-bot vendors
1,000free credits to evaluate
0charges on failed requests

Web Scraping API, Any Page, Any Vendor

Drop a Scrape action into your LlamaIndex workflow. Pass a URL, Scrapfly handles the request, stealth Chromium via Scrapium, byte-perfect Chrome TLS via Curlium, Cloudflare/Akamai/DataDome bypass via unblocker=True.

JS rendering
Cookie + session
8 anti-bot vendors
Geo proxy rotation

Why Hand-Rolled Scrapers Fail

Teams that cobble together scraping inside LlamaIndex using HTTP actions + headless browsers hit a wall within weeks.

HTTP actionblocked at TLS
Headless browser add-onfingerprint leaks
Custom code stepbreaks on every vendor update
Puppeteer serviceno anti-bot bypass
Scrapflytracked daily, 94–98%

Screenshot API

Capture any page, full-page, element, or viewport. PNG, JPEG, WebP. Ads + pop-ups auto-suppressed.

Extraction API, Structured Data via Prompt or Schema

Turn HTML into JSON inside your LlamaIndex workflow. LLM prompt or schema validation, deterministic output envelope every time.

LLM prompt extract
Schema validation
Auto extract

MCP Server

Natural-language control via Claude, ChatGPT, or LlamaIndex AI steps. No zap configuration for the model.

Workflow Recipes

Common LlamaIndex + Scrapfly patterns you can copy today.

Daily scrape → AI extract → Slack timed trigger, Scrape action, Extraction with prompt, Slack channel post
Google Sheet row → Scrape → append row row-added trigger, Scrape, Auto-extract, write back to the sheet
Monitor competitor page → Screenshot → email diff hourly trigger, Screenshot, image diff, trigger email on change

When to Pick LlamaIndex vs SDK

LlamaIndex wins on orchestration breadth. The SDK wins on tight loops + custom logic.

Under 100k req/moLlamaIndex
Needs other SaaS triggersLlamaIndex
Tight scraping loopSDK
Complex branching logicSDK
CI / cron pipelinesSDK
Python SDK


PROOF

Same API, Direct or via LlamaIndex

Under the hood, LlamaIndex actions call the same Scrapfly endpoint. Prefer code? Pick a language, every example targets the same unblocker=True path.

Set unblocker=True (formerly asp=True) and Scrapfly handles Cloudflare, Akamai, DataDome, and five other vendors. See the full bypass catalog.

from scrapfly import ScrapeConfig, ScrapflyClient, ScrapeApiResponse
client = ScrapflyClient(key="API KEY")

api_response: ScrapeApiResponse = client.scrape(
    ScrapeConfig(
        url='https://httpbin.dev/html',
        # enable the Scrapfly Unblocker
        unblocker=True
    )
)
print(api_response.result)
import { 
    ScrapflyClient, ScrapeConfig 
} from 'jsr:@scrapfly/scrapfly-sdk';

const client = new ScrapflyClient({ key: "API KEY" });
let api_result = await client.scrape(
    new ScrapeConfig({
        url: 'https://httpbin.dev/html',
        // enable the Scrapfly Unblocker
        unblocker: true,
    })
);
console.log(api_result.result);
package main

import (
	"fmt"
	"github.com/scrapfly/go-scrapfly"
)

func main() {
	client, _ := scrapfly.New("API KEY")
	result, _ := client.Scrape(&scrapfly.ScrapeConfig{
		URL: "https://httpbin.dev/html",
		// enable the Scrapfly Unblocker
		Unblocker: scrapfly.BoolPtr(true),
	})
	fmt.Println(result.Result.Content)
}
use scrapfly_sdk::{Client, ScrapeConfig};

#[tokio::main]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
    let client = Client::builder().api_key("API KEY").build()?;

    let cfg = ScrapeConfig::builder("https://httpbin.dev/html")
        // enable the Scrapfly Unblocker
        .unblocker(true)
        .build()?;

    let result = client.scrape(&cfg).await?;
    println!("{}", result.result.content);
    Ok(())
}

FAQ

Frequently Asked Questions

Is the LlamaIndex integration free?

Yes. The integration itself is free; you pay only for Scrapfly credits consumed. The free plan includes 1,000 credits with no credit card required, which is enough to evaluate the integration against your exact targets. Failed requests are not charged.


// SEE ALSO

Scrapfly plugs into every major automation platform.

Same four APIs, same unblocker=True bypass, same fingerprint coherence. Switch the orchestrator, keep the engine.