Sidecars: Web Search & Vision
Routed models do not all expose hosted web search or native image input. opencodex backfills
those capabilities with two sidecars. Both support a ChatGPT-login (forward) provider or stored
Anthropic OAuth provider; web search can additionally use stored Grok OAuth through the explicit
xai backend. Sidecar errors become bounded tool results or image markers instead of failing the
whole turn.
Additional web-search backends (explicit-only)
Section titled “Additional web-search backends (explicit-only)”Three more web-search backends exist beyond the ChatGPT and Claude paths. Each is explicit-only — it never activates from credential presence — and fails closed: a missing credential produces no sidecar plan and the request takes the normal routed path.
| Backend | Runs | Credential | Notes |
|---|---|---|---|
xai |
Grok hosted web_search (+ opt-in x_search) on api.x.ai Responses |
Stored Grok OAuth (ocx login xai) |
webSearchSidecar.xSearch enables X search with allowedXHandles/excludedXHandles (max 20, mutually exclusive) and ISO fromDate/toDate. Default model grok-4.6. |
gemini |
google_search grounding on the Antigravity transport |
Stored Antigravity OAuth with a discovered project (ocx login google-antigravity) |
Default model gemini-3.8-flash; reasoning selects the matching tier. |
exa |
Exa Search API (non-LLM result digest) | webSearchSidecar.exaApiKey |
The key is write-only through the management API (never echoed, redacted from logs). No sidecar model applies. |
Web-search sidecar
Section titled “Web-search sidecar”When Codex requests hosted web_search for a non-passthrough routed model, opencodex:
- Drops the hosted
web_searchtool and exposes a syntheticweb_search(query)function tool to the routed model instead. The original hosted-tool options are retained for the sidecar call. - Runs the routed model in a small agentic loop. When it calls
web_search, opencodex uses the selected sidecar backend: OpenAI runs hostedweb_searchwithgpt-5.6-lunaby default; Anthropic runsweb_search_20250305withclaude-sonnet-5by default. The streamed answer and citations become a tool result. xAI runs Grok hostedweb_searchwithgrok-4.6by default and, when enabled, adds hostedx_searchto the same request. - Loops until the model answers or the total real-query budget reaches
maxSearchesPerTurn(default 3), then removes the search tool and forces a final answer. Real client tools such asapply_patchor shell finalize the turn so those calls reach Codex.
When the model batches several queries into one web_search call, their results share one
8,000-character tool result. The budget is divided across the queries before anything is cut, so no
query disappears from the result. Shortened answers and unlisted sources are marked
with their original sizes, which lets the model tell trimmed results from searches that found
nothing. Structured-output turns receive the same information as one valid JSON document. If a call
asks for more queries than one result can describe, opencodex lists each query’s status without its
content and counts any queries it could not list.
Every routed-model iteration requests upstream stream: true, but by default opencodex fully
buffers semantic events internally before deciding whether to search or return the final answer.
Only the first iteration’s final headers/status and 429 key rotations are acquired eagerly. Thus
synthetic search calls and preliminary output are never exposed as client-visible model output.
Opt-in webSearchSidecar.streamRoutedModelOutput (default false) streams each iteration’s
leading text/thinking deltas live instead — the client sees output as soon as the model produces
it, exactly like the sidecar-less path. The live window closes permanently at the first tool-call
boundary, so the decision to intercept web_search stays atomic and nothing is ever delivered
twice (the terminal replay skips what already streamed). Tradeoff: text the model emits before
deciding to search — which buffered mode silently drops — becomes visible and may partially repeat
in the post-search answer. The Dashboard overview page exposes this as the Stream answers live
toggle on the web-search sidecar card (PUT /api/sidecar-settings with
webSearch.streamRoutedModelOutput).
This option also applies to adapters that manage their own turns, including Devin and Cursor. When search and image/video sidecars are both eligible, search takes priority. A first-event OAuth 429 rotates the account on the initial request and on each post-search answer request, replaying the request with the search tool and the gathered results. Cancelling a request stops subsequent searches, and retained search-loop output shares the request’s translation-buffer limit; exceeding that limit fails the response instead of starting another model iteration.
Kiro commentary is independent of this option: commentary-phase text already streams ahead of the
terminal event in buffered mode, and that bypass is unchanged — with or without
streamRoutedModelOutput, only search-decision events (tool calls and everything after the first
tool-call boundary) remain buffered for the atomic web_search decision.
The injected result is wrapped in an untrusted-data boundary, length-capped, and de-duplicated by
source URL. In structured-output turns (json_schema / json_object) it is handed over as compact
JSON instead of prose. For text-only routed models, the search model is also told to describe
relevant images in words and include their source URLs.
{ "webSearchSidecar": { "enabled": true, "backend": "anthropic", "model": "claude-sonnet-5", "reasoning": "low", "maxSearchesPerTurn": 3, "routedModelStallTimeoutMs": 200000, "timeoutMs": 200000, "streamRoutedModelOutput": false }}The explicit xAI backend uses the stored credential created by ocx login xai. Its optional
xSearch block enables X search and may restrict it to one handle list and an ISO date range:
{ "webSearchSidecar": { "backend": "xai", "model": "grok-4.6", "xSearch": { "enabled": true, "allowedXHandles": ["xai"], "fromDate": "2026-08-01", "toDate": "2026-08-21" } }}allowedXHandles and excludedXHandles are mutually exclusive and each accepts at most 20
strings. Dates use YYYY-MM-DD. Malformed management writes return 400; persisted malformed
blocks fail closed at planning time instead of silently broadening the search.
minimal reasoning is not used because the hosted backend rejects tools at that effort. A failed
search is returned to the routed model as a bounded error result, allowing it to answer from the
context it already has.
Four separate clocks apply. stallTimeoutSec is the base bridge event-stall budget.
connectTimeoutMs (default 200000) covers only DNS/TCP/TLS and final response headers.
Config-file-only webSearchSidecar.routedModelStallTimeoutMs (default 200000, integer
1..2147483647) bounds continuous raw response-byte inactivity for each routed-model iteration and
resets on every non-empty byte. webSearchSidecar.timeoutMs separately bounds one hosted search
request. The effective bridge watchdog is
max(base stall, connect timeout, routed-model stall, sidecar timeout) + 30 seconds. The routed
stall is not a total generation timeout. Failures before SSE starts return non-2xx JSON; generation
failures after response headers have started are delivered as response.failed SSE.
Vision sidecar
Section titled “Vision sidecar”Image routing is capability-aware. Before an image-bearing upstream send, opencodex resolves the selected model’s effective input modalities from runtime provider evidence, explicit operator declarations, backend/registry metadata, and generated vendor metadata. A target positively known to be text-only goes through the Vision Sidecar first; the image is described before the main call and replaced inline with text. A target positively known to support images receives the image directly. Unknown custom models keep the existing compatibility behavior rather than being guessed text-only.
For the canonical ChatGPT Codex route, opencodex uses the openai-codex metadata bundle rather than public OpenAI API metadata, so backend-specific modality differences are respected. The native Chat fast path uses the same gate and cannot bypass a known text-only verdict. Without an available sidecar plan, raw images are stripped before a proven text-only backend.
Combos advertise image input only when every member accepts images, either natively or through a
sidecar, and the combo’s imageInput setting is not disabled, so clients such as the Codex app
allow attachments instead of blocking them before the sidecar runs. When
visionSidecar.model is absent or blank, the OpenAI execution path, Dashboard, and management API
use the gpt-5.6-luna fallback. Startup still migrates an explicitly persisted legacy
gpt-5.4-mini value to gpt-5.6-luna; that migration applies to a stored value, not to an absent
model field.
The first-party DeepSeek deepseek-flash model is native multimodal (text and image) and does
not use this sidecar by default. Explicit noVisionModels or text-only declarations remain
authoritative. First-party deepseek-chat, deepseek-reasoner, and deepseek-v4-flash remain
sidecar-backed by default; Zen routes are unchanged and were not probed in this update.
- Images can come from user, developer, and tool-result messages, including Codex’s
view_image. - On the OpenAI path (ChatGPT-login passthrough), each image is sent to the configured vision model
over the Responses endpoint with the selected
reasoning.effort(lowby default), and its description replaces the image part inline. The Anthropic path uses the Messages endpoint with its own thinking-budget mapping and ignores this OpenAI-specific setting. - For native models with known capability metadata, unsupported reasoning is normalized to the highest supported rung at or below the requested level; if none exists, the lowest supported rung is used. Unknown or custom models remain permissive when reliable capability metadata is absent.
- Descriptions run with bounded concurrency (3 at a time, input order preserved). User context sent
to the describer is capped at 800 characters, and each injected description is capped at 2,000
characters. The request does not send
max_output_tokens, which the ChatGPT backend rejects. - Image URLs are validated before forwarding: data URLs must use
png/jpeg/jpg/webp/gif, and base64 data is limited to about 20 MB. Onlydata:andhttps:schemes are accepted; remotehttpsimages are fetched by the OpenAI backend, not by the proxy. noVisionModelsmatching ignores an Ollama-style:sizesuffix, so agpt-ossentry also coversgpt-oss:120b.- If description fails, the model receives a short processing-error marker. (Without an available sidecar plan, no description is attempted — the raw image is stripped, as described above.)
maxDescriptionsPerTurn(default 8) limits new descriptions per main-model turn. Cache hits and same-turn duplicates do not consume it. Successfuldata:image descriptions are cached by backend, model, detail, image bytes, and message context — plus the reasoning effort on OpenAI keys (Anthropic keys omit it, since that field is ignored there); mutablehttps:images are not cached.
The management API and Dashboard picker list models that can accept image input. When the matching
backend is available, gpt-5.6-luna (OpenAI) and claude-haiku-4-5 (Anthropic) are always offered
as baseline options. PUT /api/sidecar-settings may retain an unknown custom/ahead-of-catalog id.
An explicitly configured routed Vision Sidecar is therefore usable unless capability evidence proves
that model cannot accept images; this preserves operator-selected custom sidecars without allowing a
known text-only sidecar to receive image bytes.
{ "visionSidecar": { "enabled": true, "backend": "openai", "model": "gpt-5.6-luna", "reasoning": "medium", "maxDescriptionsPerTurn": 8, "timeoutMs": 45000 }}A model is marked text-only per provider:
{ "providers": { "ollama-cloud": { "baseUrl": "https://ollama.com/v1", "noVisionModels": ["glm-5.2", "gpt-oss", "qwen3-coder", "deepseek-v4-flash"] } }}Dashboard controls and disabling
Section titled “Dashboard controls and disabling”The Dashboard Vision sidecar card can enable or disable the sidecar, set
maxDescriptionsPerTurn, and set timeoutMs, along with the existing model,
backend, and reasoning controls. Disabling the sidecar does not delete those
settings; turning it back on keeps the previous model, backend, reasoning,
timeout, and limit.
PUT /api/sidecar-settings accepts the same fields. Partial updates leave
omitted keys unchanged. timeoutMs uses the runtime integer bounds
(1–2147483647 ms).
The web-search sidecar card carries the same control shape: the model picker’s first row is
Off. Off does two things, and the second one is the reason the row exists. OpenCodex stops
intercepting web_search, and the Codex integration writes Codex’s own
web_search = "disabled" mode into ~/.codex/config.toml — because Codex keeps declaring its
native hosted web_search tool until its own mode says otherwise, and the tool a client
advertises is the one the model reaches for. An operator who wants an MCP search server to be
the only search path needs both halves; otherwise the model keeps calling the native tool.
web_search is Codex’s key with its own value space (disabled, cached, indexed, live).
OpenCodex only ever writes disabled while the sidecar is off, and removes its marker-owned line
again once the sidecar is back on — a re-enabled sidecar whose client still had the native tool
switched off would have nothing to intercept. The write needs a managed ~/.codex/config.toml (ocx sync); the management response reports it as codexWebSearch, and both surfaces that can show it
do: the Dashboard’s web-search card warns when the write did not happen, and ocx agent sidecar web --enabled off prints whether it happened. Only a save that moves the switch triggers the write, so
the ordinary “nothing changed” answer reports not_requested and prints nothing extra. A root
web_search line the operator set by hand is replaced while the sidecar is off, since two root keys
of the same name are not valid TOML. Its exact text is recorded in the Codex journal and put back in
its place when the sidecar is switched on again — including for a line added after the journal
snapshot was taken, which ocx restore alone cannot cover. The same record is what still
recognizes our own disabled line when the Codex app has rewritten config.toml and dropped the
comment that named its owner.
You can still set enabled: false in config.json if you prefer to edit the
file directly. Anthropic-OAuth search and image description reuse the existing
Claude Code OAuth fingerprint precedent, but should be soak-tested with the
intended account and workload.
See the Configuration reference for every field.

