--provider points llmman at a model it doesn't serve itself, from
launch, from run, and from list:
export OPENROUTER_API_KEY=...
llmman providers # which providers, and is the key set
llmman list --provider openrouter # its models, and $/Mtok in and out
llmman run --provider openrouter qwen/qwen3-coder # chat with one directly
llmman launch opencode --provider openrouter --model qwen/qwen3-coder
llmman usage --since yesterday # what that session cost, per modelRequests still go through llmman serve; --provider changes where the
daemon forwards them, not who the client talks to, so local and hosted
models share one endpoint and one integration config.
The daemon also records each reply's tokens, priced at the catalog's
rates (cache, reasoning and long-context included); llmman usage sums them. See
LLMMAN_NOUSAGE in configuration.md.
The provider list comes from models.dev, the
catalog opencode uses, so a new provider needs no llmman release. It
is cached for 24 hours and a stale copy is used when the fetch fails.
When --model names a model the cached catalog lacks, the daemon re-fetches
once (at most every five minutes) before warning that the provider does not list
it. If that fetch is skipped or fails, the warning reflects the cached catalog.
All four commands read it from the daemon over
/llmman/providers, so the cache outlives any
one command and the key status reported is the daemon's.
The API key comes from the variable models.dev names for that provider,
or — when that is unset — from ~/.config/llmman/llmman.conf, keyed by
provider id:
[providers.openrouter]
api_key = "sk-or-..."or, equivalently, llmman config set providers.openrouter.api_key sk-or-...,
or from a client over HTTP: PUT /llmman/providers/openrouter/key
(what a client such as a phone app uses; see
api.md). A key set over HTTP is spent by the
running llmman serve from its next request; one set with llmman config set or a hand edit reaches a running daemon when it restarts,
like every other setting.
Either way it travels per request; it is never written into an
integration's config.
A file carrying one must be chmod 600 or its keys are ignored with a
warning; an export overrides it. See
configuration.md.
--provider needs a local llmman serve, or one reached over TLS
(LLMMAN_HOST=https://...): run and launch never send a key over
plain http to a remote LLMMAN_HOST. A daemon bound off loopback spends
its own key only for a caller that authenticated with the daemon's API
key (api.md) — and since that key takes the
Authorization header, launch's integrations then rely on the
daemon's provider key rather than carrying one; run --provider sends
its own as x-api-key. (providers and list --provider read the
catalog only and work against any daemon.)
A provider models.dev has never heard of — vLLM or llama-server on a
box down the hall, LM Studio on a laptop, a proxy in front of OpenAI —
is defined in llmman.conf by giving a [providers.<id>] a
base_url:
[providers.inferencebox]
base_url = "http://inferencebox:8000/v1"$ llmman config set providers.inferencebox.base_url http://inferencebox:8000/v1
$ llmman providers | grep inferencebox
inferencebox inferencebox - none needed -
$ llmman list --provider inferencebox # asks the box's own /models
$ llmman launch opencode --provider inferencebox --model qwen3-coderFrom there it is a provider like any other: run, list, launch and
--overflow-provider all take the id, and requests still go through
llmman serve, which forwards to the URL.
| Field | Meaning |
|---|---|
base_url |
Required to define one. An absolute http:// or https:// URL the wire's route is appended to — /chat/completions for openai, /messages for anthropic — so it usually ends in /v1. |
wire |
openai (default) or anthropic. See Wire formats. |
api_key |
Sent as the wire's credential when set. Most local servers take none, and none is sent. |
api_key_env |
An environment variable to read the key from instead; it wins over api_key, as for a catalog provider. "" clears one an earlier file named. |
name |
Display name for listings. The id when absent. |
The rules the catalog is filtered by do not apply. They vet a list
fetched from the network at runtime; a URL you wrote into your own
owner-only file needs no vetting beyond parsing. So a defined provider
may be plain http — that is the point on a LAN — and may take no key.
If it has a key and a plain-http URL, whichever process is about to
send the key warns that it crosses the network in cleartext, and sends
it.
A defined provider with a catalog id ([providers.openai] with a
base_url) replaces the catalog entry, which is how a proxy or regional
endpoint gets used without renaming the provider in every integration's
config. The catalog's model list goes with it.
Models are not listed in the file. For an openai-wire provider,
list --provider <id> and the --model check ask the endpoint's own
GET /models and take what it says; a box that is down or lacks the
route lists nothing, and the request still goes to it. An anthropic
provider has no such route and lists nothing. llmman providers shows
- in the models column for the same reason.
llmman serve reads llmman.conf once, at startup, so a provider added
while it runs needs a restart to appear. A machine that cannot reach
models.dev at all still has its defined providers.
The base_url is reported by the daemon's API and printed in warnings,
so it may not carry a user:password@; api_key is where a credential
goes. The id may not contain /.
--overflow-provider and --overflow-model pair the local --model
with a hosted one under a single name, and llmman serve picks a side
per request:
llmman launch opencode --model gemma4 --overflow-provider anthropic --overflow-model claude-sonnet-5
llmman run gemma4 --overflow-provider anthropic --overflow-model claude-sonnet-5Both halves travel as one reference,
llmman.hybrid/gemma4,anthropic/claude-sonnet-5, in the same "model"
field an ordinary name uses, so a pair works from any client on every
inference endpoint (/api/show, /api/pull and the other store
operations take a plain model name). The local half is resolved and pulled as --model always
is; the hosted half is validated and authenticates exactly as a bare
--provider model does, so the same integration rules apply. The two
cannot be combined with --provider, since the local half has to be
local.
Which side serves a request:
x-llmman-route: localorcloudon the request wins. Any other value, or the header given twice, is a400, never a guess; a blank value counts as absent.- Otherwise, size. A request larger than the local context can hold
goes to the provider. The budget is four bytes per token of the window
the local half actually loaded with — the same window
llmman launchdeclares to the integration, so what an agent is told it can send and what stays on this machine agree. Until that half is loaded, the daemon's context size (LLMMAN_CONTEXT_LENGTH) stands in.LLMMAN_HYBRID_LOCAL_BYTESsets the budget directly and is never overridden by a load,0turns the rule off. A request that declares noContent-Lengthstays local. - Otherwise, local.
Local is the default because the two mistakes are not equal: a worse
local answer is recoverable, a request sent to someone else's servers is
not. Every request logs which way it went and why; the log, not the
response's model field, is the record of the side.
The byte budget is an estimate. If a chat, completion, Responses or
Messages request it kept local is then refused by the local backend as
larger than its context, the daemon sends it to the hosted half instead,
before anything has reached the client. Without that an agent would see the local model's context error,
compact its history and stay local. A local pin is never overridden
this way, and LLMMAN_HYBRID_LOCAL_BYTES=0 disables only the size rule,
not this retry.
/v1/audio/transcriptions cannot forward to a provider, so a pair takes
its local half there whatever the body size. An unload (keep_alive: 0)
or a startup preload of a pair acts on its local half, the only one that
loads.
llmman launch with no arguments lists these and whether each is
installed:
| Name | Integration | --provider |
|---|---|---|
claude |
Claude Code | yes |
opencode |
OpenCode | yes |
codex |
OpenAI Codex CLI | yes (below) |
pi |
Pi coding agent | yes |
aider |
Aider | yes |
qwen |
Qwen Code | yes |
dsh |
DeepSeek Harness | yes |
goose |
Block goose | yes |
goose-desktop |
Block goose Desktop (requires --model) |
yes |
grok |
Grok Build (requires --model) |
no: its fetched model catalog cannot represent llmman's hosted-provider routing reference |
docker-agent |
Docker Agent (requires --model) |
yes |
hermes |
Hermes Agent | yes, but the daemon holds the key (below) |
agy |
Antigravity CLI (requires --model) |
yes |
gemini |
Gemini CLI | no: llmman cannot confirm the key would come here rather than go to Google |
cline |
Cline (requires --model) |
yes |
kimi |
Kimi Code CLI | no: it picks its own model rather than taking llmman's |
copilot |
GitHub Copilot CLI (gh) |
no: it has no way to send a key |
openclaw |
OpenClaw | no: it only takes a model during first-run onboarding |
Any extra arguments after -- are forwarded to the integration's own CLI.
hermes is configured through a file on disk, so it can't carry a key
per request; llmman serve needs one of its own, spent only for a
loopback daemon and never for a cross-site browser request. On a shared
machine prefer an integration that sends its own key.
codex speaks only OpenAI's Responses API, which most providers lack
(mistral 404s it, opencode 500s it for non-OpenAI models). The
daemon tries the provider first and, on a 404/405/501 or 5xx, translates
the request to a chat completion and the reply back, tool calls included.
Providers that have the API (openai, groq, openrouter) are used
natively; any other 4xx is relayed as-is.
Thinking depth is set from inside the integration and reaches the model
as reasoning_effort: llama-server reads it natively (none turns
thinking off; a level goes to the chat template), a provider gets it in
its own form (see wire formats). Nothing selected leaves
the model's default.
opencode: variants read off the model's chat template (whatllmman showlists asthinking), cycled withvariant_cycle(ctrl+t) or/variants:none, eachreasoning_effortlevel the template takes (Qwen3.8:low,medium,xhigh), orthinkingfor a template with only anenable_thinkingswitch (Gemma 4, Qwen3.5). A provider's model gets the levels models.dev'sreasoning_optionslist, as opencode does (Claude Opus 5.5:lowtomax; an Anthropicbudget_tokensmodel:high,max); none where models.dev says it does not reason; elsenone,low,medium,high, plus the--variantasked for.claude: Claude Code's/effort <low|medium|high|xhigh|max>, sent as spelled; a level the template rejects is a 400.codex:model_reasoning_effort, e.g.-- -c model_reasoning_effort=high; its/modelpicker lists only OpenAI's catalog.
--variant, as on opencode's run, picks one of those up front:
llmman run qwen3.8 --variant xhigh sends it as is; llmman launch
first checks the model has it, then hands it over in the integration's
own form, still changeable from inside. A model with no known levels,
such as one models.dev does not list yet, takes any effort level.
| Integration | As | Takes |
|---|---|---|
opencode |
the model's default options and its saved variant | every variant |
claude, copilot |
--effort |
low to max |
codex |
-c model_reasoning_effort= |
all |
pi, omp |
--thinking |
none, minimal to high (omp: xhigh) |
cline |
--thinking |
none, low to high |
aider |
--reasoning-effort |
all |
hermes |
--reasoning (not with --tui) |
all |
grok |
--effort |
all |
qwen |
its model entry | all but minimal |
dsh |
its route | all |
thinking goes to all but opencode as medium, which llmman serves as
thinking on. kimi, openclaw, gemini, agy, goose,
goose-desktop and docker-agent cannot carry a variant here, so they
refuse one.
Each provider is spoken to in one of two wire formats, reported as
wire by /llmman/providers:
openai: OpenAI Chat Completions withAuthorization: Bearer <key>. Every@ai-sdk/openai-compatibleprovider, plus the hand-checked endpoints foropenai,google,groq,mistraland the rest, and the default for a provider defined inllmman.conf.anthropic: the Anthropic Messages API withx-api-key: <key>.anthropicitself. Other Messages-compatible endpoints are not offered from the catalog, since their auth scheme varies and has not been checked; one you know takesx-api-keycan be defined withwire = "anthropic".
A provider that takes no key gets no credential header at all, not an empty one.
Anthropic is never reached through its OpenAI-compatibility shim. What a
request becomes on the way to a wire: anthropic provider depends on
the surface it arrived on:
| Arrived on | Sent as |
|---|---|
/v1/messages (Claude Code) |
The same request, relayed intact: cache breakpoints, thinking, tools and anthropic-beta headers included. Only model is rewritten. |
/v1/chat/completions (OpenCode, Aider, Qwen Code, Hermes), /api/chat, /api/generate |
A Messages request, and the reply back as chat-completion chunks: system turns to system, tool calls to tool_use/tool_result, reasoning_effort to thinking (see below), thinking back as reasoning_content. |
/v1/responses (Codex) |
The Responses bridge above, then the same translation. The provider is not probed for /v1/responses. |
max_tokens is required by the Messages API; a translated request
without one gets the model's limit.output from the catalog, or 4096
(a relayed /v1/messages request is the client's own to complete). /v1/completions,
/v1/embeddings, /api/embed, /api/embeddings and
/v1/responses/input_tokens have no Messages equivalent and are refused
with a 501.
reasoning_effort takes the form the model accepts, by the version in
a Claude's name. From Claude 4.6 (Sonnet 5, Opus 4.7, an unversioned
preview) it is adaptive thinking with the level as output_config.effort
(minimal as low); none turns thinking off, or on Fable and Mythos,
which always think, is low. Through Claude 4.5, and on another
vendor's Messages endpoint, it is a budget spent from max_tokens, left
off for the continuation of a tool call or a forced tool, which want a
signed thinking block no OpenAI client can hand back.
The translation also does what the API needs that an OpenAI client would
not know to: prompt caching is on (breakpoints on the last tool, system
block and user block), a response_format JSON schema becomes a forced
tool whose arguments are returned as the reply, tools used earlier in
the history are declared back when the client offers none, and an
unanswered tool call gets a placeholder result.
Enable [managed] in configuration.md
to forward native Codex and Claude requests on the daemon's existing TLS
listener. The caller owns login, refresh, account selection, authorization,
and network policy. llmman never persists these request-local OAuth tokens or
falls back to stored provider keys, environment credentials, or peers.
Every forwarding request must come from a loopback connection and supply all of:
X-Api-Key: <configured daemon API key>
Authorization: Bearer <current provider OAuth access token>
X-LLMMan-Upstream-IP: <policy-authorized public IP for the provider>
The daemon key is checked separately and removed before forwarding. An OAuth
bearer alone cannot authenticate to the daemon. Peer admission uses the actual
connection address, never Forwarded or X-Forwarded-For. Neither disabling
daemon authentication nor configuring an HTTP listener is allowed in managed mode.
| Managed operation | Fixed upstream | Required model prefix |
|---|---|---|
POST /api/codex/responses |
https://chatgpt.com/backend-api/codex/responses |
llmman.provider/openai/ |
POST /api/codex/responses/compact |
https://chatgpt.com/backend-api/codex/responses/compact |
llmman.provider/openai/ |
POST /api/anthropic/messages |
https://api.anthropic.com/v1/messages |
llmman.provider/anthropic/ |
POST /api/anthropic/messages/count_tokens |
https://api.anthropic.com/v1/messages/count_tokens |
llmman.provider/anthropic/ |
The route selects the provider; no auth-profile header or account-header
heuristic is needed. ChatGPT-Account-ID is optional on Codex requests and is
never sent to Anthropic. The exact nonempty model suffix is forwarded after
removing the table's prefix. Other native payload fields are preserved; invalid
credentials, mismatched model prefixes, and duplicate top-level JSON fields
are rejected before contacting a provider. Request bodies are bounded to 32 MiB.
Encoded request bodies are rejected with 415; absent or case-insensitive
identity content coding is accepted and removed before JSON is reserialized.
The supervisor resolves the fixed provider hostname, authorizes a particular
IP against its network policy, and supplies that numeric address in
X-LLMMan-Upstream-IP. llmman pins the connection to it while verifying the
provider's TLS hostname. There is no DNS fallback or caller-controlled URL.
Special-use addresses, including private networks, loopback, and IPv6 6to4,
are rejected. Redirects and environment proxy/CA overrides are not used.
Provider extension headers pass through in both directions, including repeated
values. Hop-by-hop headers, fields named by Connection, Host, internal
X-LLMMan-* headers, daemon/alternate credentials, and cookies are removed.
Content-Length is regenerated after body changes. The provider bearer and
Codex account ID are set explicitly. Codex originator and openai-beta are
forwarded only when supplied; llmman does not invent a client identity.
For Claude, llmman merges oauth-2025-04-20 into anthropic-beta and defaults
anthropic-version to 2023-06-01 when absent. The caller supplies any required
native system content or metadata.
Successful response bodies stream with backpressure and cancellation on
disconnect. Upstream connection establishment has a 30-second timeout; reads
have no deadline unless managed.read_timeout_seconds is configured. That
optional timeout applies to response-header waits and individual stalled reads,
not total generation duration. A continuously active stream can outlive this
limit; a quiet generation is interrupted when the limit expires. Choose a
value appropriate for the workload.
Provider error statuses are preserved with bearer/account values redacted from error bodies and response headers. JSON error bodies are decoded before redacting string values and keys, so escaped credentials are removed too. Error bodies above 1 MiB, interrupted reads, redirects, and encoded error bodies that cannot be safely redacted produce 502. llmman requests uncompressed upstream responses. Transport diagnostics are redacted and available with debug logging. A failure after streaming headers have been sent terminates the stream instead of changing its status.
GET /api/version confirms daemon liveness. Before supplying provider tokens,
call GET /api/managed/capabilities over a trusted TLS connection with the
same daemon API key, and require the corresponding capability:
{"capabilities":["codex-oauth-forwarding-v2","claude-oauth-forwarding-v2"]}These routes bypass prompt-history recording and ordinary provider routing. They are disabled by default. The earlier draft's second listener, JSON config, ready file, and managed-key/auth-profile headers have been removed; callers of that draft must migrate configuration, paths, authentication, and capability checks together. A supervisor must bind readiness to the child it launched and invalidate it on exit; a version response alone is not a forwarding handshake.