Skip to content

[FEATURE] WebLLM model provider (on-device inference via WebGPU) #2481

Description

@jsamuel1

Problem Statement

The Strands TypeScript SDK does not ship a first-class model provider for WebLLM. WebLLM runs quantized LLMs entirely in the browser via WebGPU, with model weights cached in IndexedDB/CacheStorage after the first download. Without a dedicated provider, users who want on-device, offline-capable agents have to wire WebLLM up themselves.

Use Case

  • Building browser-based agent apps that run fully offline after the initial model download
  • Privacy-sensitive deployments where prompts/responses must not leave the user's device
  • Eliminating per-call API costs by running inference on user hardware
  • Demo/educational apps that work without any cloud credentials

Proposed Solution

A WebLLMModel provider under @strands-agents/sdk/models/webllm that:

  • Wraps @mlc-ai/web-llm (MLCEngine / CreateMLCEngine)
  • Implements the standard Model interface (streaming, tool use, stop reasons, usage)
  • Exposes helpers for cache management: downloadWebLLMModel, isWebLLMModelCached, deleteWebLLMModel, listWebLLMModels
  • Surfaces InitProgressReport via an onProgress callback so apps can render download progress
  • Guards against non-browser environments (WebLLM requires WebGPU)

Workaround

Today you can use VercelModel with the community webllm-ai-provider package (Vercel AI SDK v2 provider spec):

import { Agent } from '@strands-agents/sdk'
import { VercelModel } from '@strands-agents/sdk/models/vercel'
import { webllm } from 'webllm-ai-provider'

const agent = new Agent({
  model: new VercelModel({ model: webllm() }),
})

This works but has caveats:

  • webllm-ai-provider is a third-party package at 0.0.1 with a single maintainer and ~2 weekly downloads
  • No built-in cache management API (download-before-use, evict, list-cached)
  • Progress reporting goes through the provider rather than a first-class onProgress hook
  • WebGPU/browser-only environment checks are the user's responsibility

Alternatives

  1. Keep recommending the VercelModel + webllm-ai-provider workaround in docs only
  2. Consume @mlc-ai/web-llm directly via a custom Model subclass in user code
  3. Ship a dedicated WebLLMModel provider (proposed)

Additional Context

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Fields

    Language

    TypeScript

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions