Skip to content

Latest commit

 

History

History
384 lines (297 loc) · 14.8 KB

File metadata and controls

384 lines (297 loc) · 14.8 KB

langchain-cloudflare

This package contains the LangChain integration with CloudflareWorkersAI

Installation

pip install -U langchain-cloudflare

And you should configure credentials by setting the following environment variables:

  • CF_ACCOUNT_ID

AND

  • CF_API_TOKEN (if using a single token scoped for all services)

OR (if using separately scoped tokens)

  • CF_AI_API_TOKEN (CloudflareWorkersAI, CloudflareWorkersAIEmbeddings, CloudflareBrowserRunLoader, CloudflareBrowserRunTool)
  • CF_AI_SEARCH_API_TOKEN (CloudflareAISearchRetriever)
  • CF_VECTORIZE_API_TOKEN (CloudflareVectorize)
  • CF_D1_API_TOKEN (CloudflareVectorize)
  • CF_D1_DATABASE_ID (CloudflareVectorize)

Browser Run requires the Browser Rendering – Edit permission on your API token. See Browser Run setup.

Chat Models

ChatCloudflareWorkersAI class exposes chat models from CloudflareWorkersAI.

from langchain_cloudflare.chat_models import ChatCloudflareWorkersAI

llm = ChatCloudflareWorkersAI()
llm.invoke("Sing a ballad of LangChain.")

REST endpoint format

By default, ChatCloudflareWorkersAI uses the native Workers AI run endpoint:

llm = ChatCloudflareWorkersAI(
    model="@cf/moonshotai/kimi-k2.6",
    endpoint_format="workers_ai",  # default
)

For REST calls that need Cloudflare's OpenAI-compatible chat completions API, set endpoint_format="openai_compatible":

llm = ChatCloudflareWorkersAI(
    model="@cf/moonshotai/kimi-k2.6",
    endpoint_format="openai_compatible",
)

When ai_gateway is configured, OpenAI-compatible mode routes through the Workers AI chat completions path on AI Gateway. This option is REST-only; Worker bindings use env.AI.run() and do not expose a chat completions route.

For chat dynamic routes (model="dynamic/<route name>"), message response_metadata["model_name"] identifies the model reported by the provider. response_metadata["requested_model"] retains the requested route. If the provider does not report a model, model_name is None. Each message preserves its model identity when a batch uses different fallback models.

Decision Models (Clef)

CloudflareWorkersAIDecisionModel supports @cf/cloudflare/clef and @cf/cloudflare/clef-flash. They evaluate text or JSON state against typed questions and return probabilities. Question types are noul (yes/no probability), choice (an option with its probability distribution), and score (a probability-weighted ordered rubric).

from langchain_cloudflare import CloudflareWorkersAIDecisionModel

model = CloudflareWorkersAIDecisionModel(model_name="@cf/cloudflare/clef-flash")
questions = {
    "urgent": {"type": "noul", "instructions": "Is this request urgent?"},
    "team": {
        "type": "choice",
        "instructions": "Which team should handle this request?",
        "criteria": {
            "billing": "Invoices and refunds",
            "technical": "Outages and errors",
        },
    },
    "severity": {
        "type": "score",
        "instructions": "How severe is the impact?",
        "criteria": ["No impact", "Minor", "Major", "Critical"],
    },
}
result = model.evaluate(
    state="Checkout is failing for every customer.", questions=questions
)
print(result["answers"])  # Includes probabilities, confidence, and scores.

# Async REST uses the same input and result shape.
result = await model.aevaluate(
    state={"service": "checkout", "status": "down"}, questions=questions
)

# In a Python Worker, authentication comes from the existing AI binding.
model = CloudflareWorkersAIDecisionModel(binding=self.env.AI)
result = await model.aevaluate(state="Checkout is down.", questions=questions)

REST credentials use CF_ACCOUNT_ID and CF_AI_API_TOKEN. Configure ai_gateway to use an AI Gateway. reject_if_busy=True passes live Clef Worker binding tests; Clef's REST validator currently rejects the documented options field with HTTP 422. Clef and Clef Flash are not currently supported through dynamic routing. Results preserve model, answers, and usage fields. Pass optional images as embedded PNG/JPEG/WebP data URLs or {"content_type": "image/png", "base64": "..."} objects (up to four images). The Worker example exposes this operation at /decision.

Embeddings

CloudflareWorkersAIEmbeddings class exposes embeddings from CloudflareWorkersAI.

from langchain_cloudflare.embeddings import CloudflareWorkersAIEmbeddings

embeddings = CloudflareWorkersAIEmbeddings(model_name="@cf/baai/bge-base-en-v1.5")
embeddings.embed_query("What is the meaning of life?")

VectorStores

CloudflareVectorize class exposes vectorstores from Cloudflare Vectorize.

from langchain_cloudflare.vectorstores import CloudflareVectorize

vst = CloudflareVectorize(embedding=embeddings)
vst.create_index(index_name="my-cool-vectorstore")

Retrievers

CloudflareAISearchRetriever exposes Cloudflare AI Search (the managed retrieval / RAG service, fka AutoRAG) as a LangChain retriever.

Prerequisites

  • An AI Search instance with content. The retriever searches an existing instance, so create one and add your data first — via the dashboard, Wrangler, or the Python SDK.
  • Credentials, read from the environment:
    • CF_ACCOUNT_ID
    • CF_AI_SEARCH_API_TOKEN — an AI Search:Run token (falls back to CF_API_TOKEN)
    • CF_AI_SEARCH_INSTANCE_NAME — or pass instance_name=

Usage

from langchain_cloudflare import CloudflareAISearchRetriever

retriever = CloudflareAISearchRetriever(instance_name="my-instance")
docs = retriever.invoke("How do I configure Workers AI?")

Inside a Python Worker, pass the dedicated ai_search binding instead of REST credentials (async only):

retriever = CloudflareAISearchRetriever(binding=env.MY_SEARCH)
docs = await retriever.ainvoke("How do I configure Workers AI?")

The constructor exposes AI Search's retrieval options (hybrid search, metadata filters, reranking, query rewriting, …) as parameters, plus an ai_search_options parameter for passing any AI Search option that doesn't have its own parameter. As a standard BaseRetriever it plugs into RAG chains and becomes an agent tool via create_retriever_tool. For multi-tenant setups, give each tenant its own instance and point a retriever at that instance.

Browser Run: REST vs. Worker Binding Parity

Browser Run has two distinct APIs: Quick Actions (single request/response calls -- markdown extraction, screenshots, structured extraction, etc.) and full browser sessions (stateful, multi-step control via CDP/Puppeteer/Playwright/Stagehand -- click, type, navigate across pages). This library only implements Quick Actions; full sessions are JS/npm-only (@cloudflare/puppeteer, Playwright) with no Python equivalent, so they're not reachable from a Python Worker at all, REST or binding.

Every Quick Action is reachable through this library, split across CloudflareBrowserRunLoader (document ingestion) and CloudflareBrowserRunTool (agent actions) — see their sections below for details. crawl and browser="kitesurf" are REST-only; every other Quick Action works over both REST and the binding parameter, verified live against the real API and against a real Python Worker:

Mode Class REST Binding (quickAction())
markdown Loader, Tool ✅ ✅
content Loader ✅ ✅
scrape Loader ✅ ✅
crawl Loader ✅ ❌ async job with polling, no quickAction() equivalent
json Tool ✅ ✅
links Tool ✅ ✅
screenshot Tool ✅ ✅
pdf Tool ✅ ✅
snapshot Tool ✅ ✅
accessibility_tree Tool ✅ ✅
browser="kitesurf" Loader, Tool ✅ ❌ URL query param, no binding equivalent

The binding path is async-only on both classes (aload()/ainvoke(), not load()/invoke()) — calling the sync methods with binding set raises NotImplementedError.

Browser Run (Document Loader)

CloudflareBrowserRunLoader loads web pages as LangChain Document objects using Cloudflare Browser Run (formerly Browser Rendering). It renders JavaScript-heavy pages on Cloudflare's global network and returns clean content via a REST API or, inside a Python Worker, the browser binding.

from langchain_cloudflare import CloudflareBrowserRunLoader

# Single page -> markdown
loader = CloudflareBrowserRunLoader(
    urls=["https://developers.cloudflare.com/workers-ai/"],
    mode="markdown",
)
docs = loader.load()

# Multi-page crawl -> knowledge base (REST-only; async job with polling)
loader = CloudflareBrowserRunLoader(
    urls=["https://developers.cloudflare.com/cloudflare-one/"],
    mode="crawl",
    crawl_limit=50,
    crawl_depth=2,
    crawl_options={"source": "sitemaps"},  # any other /crawl body option
)
docs = loader.load()

# Scrape specific elements with CSS selectors
loader = CloudflareBrowserRunLoader(
    urls=["https://example.com/pricing"],
    mode="scrape",
    elements=[{"selector": "h1"}, {"selector": ".plan-card"}],
)
docs = loader.load()  # one Document per matched selector group

# Async support
docs = await loader.aload()

Supported modes:

Mode Endpoint Description
markdown /markdown Clean markdown from any page
crawl /crawl Multi-page crawl with async polling (REST-only)
scrape /scrape CSS selector-based element extraction
content /content Raw rendered HTML

Inside a Python Worker, pass the browser binding instead of REST credentials (async only — use aload()/alazy_load(), not load()/lazy_load()):

loader = CloudflareBrowserRunLoader(
    urls=["https://example.com"], mode="markdown", binding=env.BROWSER
)
docs = await loader.aload()

Pass browser="kitesurf" to use Cloudflare's stateless, agent-optimized browser runtime instead of full Chromium (REST-only — not reachable via the quickAction() binding, since it's a URL query parameter with no equivalent in the binding's params object):

loader = CloudflareBrowserRunLoader(
    urls=["https://example.com"], mode="markdown", browser="kitesurf"
)

Browser Run (Agent Tool)

CloudflareBrowserRunTool gives LangGraph agents the ability to interact with the live web.

from langchain_cloudflare import CloudflareBrowserRunTool

# Read any page as markdown
tool = CloudflareBrowserRunTool(mode="markdown")
content = tool.invoke({"url": "https://example.com"})

# AI-powered structured data extraction
tool = CloudflareBrowserRunTool(
    mode="json",
    json_prompt="Extract the company name, pricing plans, and key features.",
)
data = tool.invoke({"url": "https://www.cloudflare.com/plans/"})

# Combined-format snapshot (markdown + screenshot in one call)
tool = CloudflareBrowserRunTool(
    mode="snapshot", snapshot_formats=["markdown", "screenshot"]
)
snapshot = tool.invoke({"url": "https://example.com"})

# Accessibility tree (roles, names, states, hierarchy)
tool = CloudflareBrowserRunTool(mode="accessibility_tree")
tree = tool.invoke({"url": "https://example.com"})

# Use multiple tools in a LangGraph agent
from langgraph.prebuilt import ToolNode

tools = [
    CloudflareBrowserRunTool(mode="markdown"),
    CloudflareBrowserRunTool(mode="json", json_prompt="Extract key facts."),
    CloudflareBrowserRunTool(mode="links"),
]
tool_node = ToolNode(
    tools
)  # each tool auto-named: cloudflare_browser_run_markdown, etc.

Supported modes:

Mode Endpoint Description
markdown /markdown Read any webpage as markdown
json /json AI-powered structured data extraction
links /links Discover all links on a page
screenshot /screenshot Capture screenshot (base64 PNG)
pdf /pdf Generate PDF (base64)
snapshot /snapshot Multiple page formats in one call
accessibility_tree /accessibilityTree Accessibility tree as JSON

Inside a Python Worker, pass the browser binding instead of REST credentials (async only — use ainvoke(), not invoke()). This calls Browser Run's quickAction() RPC method instead of the REST API:

tool = CloudflareBrowserRunTool(mode="markdown", binding=env.BROWSER)
result = await tool.ainvoke({"url": "https://example.com"})

The browser binding requires a compatibility_date of 2026-03-24 or later and, in local development, "remote": true (quickAction() isn't supported in local simulation):

// wrangler.jsonc
{
  "compatibility_date": "2026-03-24",
  "browser": { "binding": "BROWSER", "remote": true }
}

Release Notes

v0.1.1 (2025-04-08)

  • Added ChatCloudflareWorkersAI integration
  • Added CloudflareWorkersAIEmbeddings support
  • Added CloudflareVectorize integration

v0.1.3 (2025-04-10)

  • Added AI Gateway support for CloudflareWorkersAIEmbeddings
  • Added Async support for CloudflareWorkersAIEmbeddings

v0.1.4 (2025-04-14)

  • Added support for additional model parameters as explicit class attributes for ChatCloudflareWorkersAI

v0.1.6 (2025-05-01)

  • Added Standalone D1 Metadata Filtering Methods
  • Update Docs for more clarity around D1 Table/Vectorize Index Names

v0.1.8 (2025-05-11)

  • Added support for environmental variables (embeddings, vectorstores)