Skip to content

Commit 34fd3cb

Browse files
committed
feat(inference): inference interception and routing (!38)
Closes NVIDIA#67 ## Summary Implements transparent inference interception and routing for sandboxed AI agents. The sandbox proxy intercepts outbound AI SDK calls (OpenAI, Anthropic) and reroutes them through the gateway to policy-controlled backends — enabling organizations to redirect inference traffic to local or self-hosted models without modifying agent code. **Decision model** — a tri-state OPA evaluation for every CONNECT request: 1. Binary + endpoint explicitly allowed in `network_policies` → **allow** (pass through) 2. Not explicitly allowed + `inference.allowed_routes` configured → **inspect for inference** (TLS intercept, detect API patterns, route through gateway) 3. Otherwise → **deny** No endpoint declarations or binary lists needed for inference routing. Just configure `inference.allowed_routes`. ## Key Changes ### Sandbox (interception) - **OPA policy**: New `network_action` Rego rule with three outcomes (`allow`, `inspect_for_inference`, `deny`). New `NetworkAction` enum replaces `PolicyDecision.allowed` bool for the proxy's main decision path. - **Proxy**: New `InspectForInference` path — TLS-terminates client, parses HTTP, detects inference API patterns (`POST /v1/chat/completions`, `/v1/completions`, `/v1/messages`), strips auth headers, forwards via gRPC. - **New module**: `l7/inference.rs` — `InferenceApiPattern`, `detect_inference_pattern()`, HTTP request/response parsing. - **gRPC client**: New `proxy_inference()` for sandbox→gateway forwarding. - **Sandbox init**: Creates OPA engine when inference is configured, even without `network_policies`. ### Gateway (dispatch) - **InferenceService**: `ProxyInference` RPC loads sandbox policy, resolves allowed routes, dispatches to router. Full CRUD for inference routes. - **Proto**: `InferenceRoute`, `InferenceRouteSpec`, `ProxyInferenceRequest/Response`, Inference gRPC service. ### Router (backend proxying) - **New crate**: `navigator-router` with `Router`, `proxy_with_candidates()`, protocol-based route selection, backend HTTP proxying with auth header rewriting. - **Mock support** for testing (`mock://` scheme). ### CLI - `nav inference create/update/delete/list` commands for route management. ### Python SDK - Updated protobuf bindings. Removed old `inference.py` client (replaced by transparent interception). ### Documentation - New `architecture/inference-routing.md` — end-to-end system documentation. - Updated `architecture/sandbox.md` — proxy, OPA, and source index sections. - Updated `architecture/README.md` — new subsystem overview and diagram. ## Addendum: Chunked Transfer Compatibility This branch now also fixes intercepted SDK requests that send chunked request bodies: - `inspect_for_inference` now accepts `Transfer-Encoding: chunked` and decodes chunked request bodies before forwarding to the gateway - Removed the prior `411 Length Required` response for chunked intercepted requests - Added request/response header sanitization for framing and hop-by-hop headers (`content-length`, `transfer-encoding`, `connection`, etc.) to keep forwarded requests and returned responses valid - Added unit tests for chunked parsing and header sanitization Note: this improves compatibility for streaming-style SDK request patterns; true token-by-token passthrough response streaming is still a separate follow-up. ## Minimal Policy for Inference Routing ```yaml inference: allowed_routes: - local ``` Any outgoing connection from a binary not explicitly allowed in `network_policies` will be intercepted and checked for inference API patterns. ## Test Plan - [x] `cargo test --workspace` — all tests pass - [x] `mise run pre-commit` — all checks pass - [x] E2E: OpenAI chat completions routed through gateway - [x] E2E: Anthropic messages routed through gateway - [x] E2E: Python OpenAI SDK from sandbox (`examples/inference/inference.py`) - [x] E2E test: `e2e/python/test_inference_routing.py`
1 parent 2f808ea commit 34fd3cb

45 files changed

Lines changed: 3406 additions & 1344 deletions

Some content is hidden

Large Commits have some content hidden by default. Use the searchbox below for content that may be hidden.

‎.claude/agent-memory/arch-doc-writer/MEMORY.md‎

Lines changed: 15 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -7,7 +7,7 @@
77
- Sandbox entry: `crates/navigator-sandbox/src/lib.rs` (`run_sandbox()`)
88
- OPA engine: `crates/navigator-sandbox/src/opa.rs` (single file, not a directory)
99
- Identity cache: `crates/navigator-sandbox/src/identity.rs` (SHA256 TOFU, uses Mutex<HashMap> NOT DashMap)
10-
- L7 inspection: `crates/navigator-sandbox/src/l7/` (mod.rs, tls.rs, relay.rs, rest.rs, provider.rs)
10+
- L7 inspection: `crates/navigator-sandbox/src/l7/` (mod.rs, tls.rs, relay.rs, rest.rs, provider.rs, inference.rs)
1111
- Proxy: `crates/navigator-sandbox/src/proxy.rs`
1212
- Server multiplex: `crates/navigator-server/src/multiplex.rs`
1313
- SSH tunnel: `crates/navigator-server/src/ssh_tunnel.rs`
@@ -18,7 +18,7 @@
1818

1919
## Architecture Docs
2020
- Files renamed from numbered prefix format to descriptive names (e.g., `2 - server-architecture.md` -> `gateway-architecture.md`)
21-
- Current files: README.md, sandbox-providers.md, cluster-single-node.md, build-containers.md, sandbox-connect.md, sandbox.md, security-policy.md, gateway.md
21+
- Current files: README.md, sandbox-providers.md, cluster-single-node.md, build-containers.md, sandbox-connect.md, sandbox.md, security-policy.md, gateway.md, inference-routing.md, inference-routing-debug-handoff.md
2222
- Cross-references use plain filenames: `[text](gateway.md)`
2323
- Naming convention: "gateway" in prose for the control plane component; code identifiers like `navigator-server` stay unchanged
2424

@@ -101,6 +101,19 @@
101101
- DNS failure also rejects the connection
102102
- Non-CP connections use pre-resolved addrs: `TcpStream::connect(addrs.as_slice())`
103103

104+
## Inference Routing Details
105+
- Three-action OPA model: Allow, InspectForInference, Deny (see `NetworkAction` in opa.rs)
106+
- InspectForInference triggers: no explicit network_policy match AND inference.allowed_routes is non-empty in policy
107+
- Proxy handler: `handle_inference_interception()` in proxy.rs -- TLS-terminates, parses HTTP, pattern-matches, calls gateway
108+
- gRPC client: `proxy_inference()` in `crates/navigator-sandbox/src/grpc_client.rs`
109+
- Default inference patterns: POST /v1/chat/completions (openai_chat_completions), POST /v1/completions (openai_completions), POST /v1/messages (anthropic_messages)
110+
- Pattern detection: exact path match after stripping query string, case-insensitive method
111+
- Gateway route resolution: loads all InferenceRoute objects, filters by enabled + routing_hint in allowed_routes, then router selects by protocol match
112+
- Router: `navigator-router` crate, `proxy_with_candidates()` finds first route whose `protocols` list contains the source_protocol
113+
- InferenceRouteSpec proto fields: routing_hint, base_url, protocols (repeated string), api_key, model_id, enabled
114+
- Auth header stripping: proxy removes `Authorization` header before forwarding; gateway/router adds route's api_key as Bearer token
115+
- Non-inference requests on intercepted connections get 403 JSON response
116+
104117
## Naming Conventions
105118
- The project name "Navigator" appears in code but docs should use generic terms per user preference
106119
- CLI binary: `navigator` (aliased as `nav` in dev via mise)

‎Cargo.lock‎

Lines changed: 1 addition & 0 deletions
Some generated files are not rendered by default. Learn more about customizing how changed files appear on GitHub.

‎architecture/README.md‎

Lines changed: 52 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -33,6 +33,7 @@ flowchart TB
3333
subgraph EXT["External Services"]
3434
HOSTS["Allowed Hosts (github.com, api.anthropic.com, ...)"]
3535
CREDS["Provider APIs (Claude, GitHub, GitLab, ...)"]
36+
BACKEND["Inference Backends (LM Studio, vLLM, ...)"]
3637
end
3738
3839
CLI -- "gRPC / HTTPS" --> SERVER
@@ -44,6 +45,8 @@ flowchart TB
4445
CHILD -- "All network traffic" --> PROXY
4546
PROXY -- "Evaluate request" --> OPA
4647
PROXY -- "Allowed traffic only" --> HOSTS
48+
PROXY -- "Inference reroute (gRPC)" --> SERVER
49+
SERVER -- "Proxied inference" --> BACKEND
4750
SERVER -. "Store / retrieve credentials" .-> CREDS
4851
```
4952

@@ -80,6 +83,7 @@ When the agent (or any tool running inside the sandbox) tries to connect to a re
8083
3. **Evaluates the request against policy** using the OPA engine. The policy can allow or deny connections based on the destination hostname, port, and the identity of the requesting program.
8184
4. **Rejects connections to internal IP addresses** as a defense against SSRF (Server-Side Request Forgery). Even if the policy allows a hostname, the proxy resolves DNS before connecting and blocks any result that points to a private network address (e.g., cloud metadata endpoints, localhost, or RFC 1918 ranges). This prevents an attacker from redirecting an allowed hostname to internal infrastructure.
8285
5. **Performs protocol-aware inspection (L7)** for configured endpoints. The proxy can terminate TLS, inspect the underlying HTTP traffic, and enforce rules on individual API requests -- not just connection-level allow/deny. This operates in either audit mode (log violations but allow traffic) or enforce mode (block violations).
86+
6. **Intercepts inference API calls** when the sandbox has inference routing configured. Connections that don't match any explicit network policy but have inference routes available are TLS-terminated and inspected. Known inference API patterns (OpenAI, Anthropic) are detected and rerouted through the gateway to the configured backend, while non-inference requests are denied.
8387

8488
The proxy generates an ephemeral certificate authority at startup and injects it into the sandbox's trust store. This allows it to transparently inspect HTTPS traffic when L7 inspection is configured for an endpoint.
8589

@@ -158,6 +162,51 @@ This approach means users configure credentials once, and every sandbox that nee
158162

159163
For more detail, see [Providers](sandbox-providers.md).
160164

165+
### Inference Routing
166+
167+
The inference routing system transparently intercepts AI inference API calls from sandboxed agents and reroutes them through the gateway to policy-controlled backends. This enables organizations to redirect inference traffic to local or self-hosted models without modifying the agent's code.
168+
169+
**How it works end-to-end:**
170+
171+
1. The sandbox policy includes an `inference.allowed_routes` list (e.g., `["local"]`).
172+
2. When the agent makes an HTTPS request to any endpoint (e.g., `api.openai.com`), the proxy evaluates it:
173+
- If the endpoint + binary is explicitly allowed in `network_policies` -- pass through directly.
174+
- If no policy match but inference routes are configured -- **intercept** (OPA returns the `inspect_for_inference` action).
175+
- Otherwise -- deny.
176+
3. For intercepted connections, the proxy:
177+
- TLS-terminates the client connection using the sandbox's ephemeral CA.
178+
- Parses the HTTP request.
179+
- Detects known inference API patterns (e.g., `POST /v1/chat/completions` for OpenAI, `POST /v1/messages` for Anthropic).
180+
- Strips authorization headers and forwards the request to the gateway via gRPC (`ProxyInference` RPC).
181+
4. The gateway's inference service:
182+
- Loads the sandbox's policy to get `allowed_routes`.
183+
- Finds enabled inference routes whose `routing_hint` matches the allowed list.
184+
- Selects a compatible route by matching the source protocol (e.g., `openai_chat_completions`).
185+
- Forwards the request to the route's backend URL, rewriting the authorization header with the route's API key.
186+
5. The response flows back through the gateway to the proxy to the agent -- the agent sees a normal HTTP response as if it came from the original API.
187+
188+
**Key design properties:**
189+
190+
- Agents need zero code changes -- standard OpenAI/Anthropic SDK calls work transparently.
191+
- The sandbox never sees the real API key for the backend -- credential isolation is maintained.
192+
- Policy controls which routes a sandbox can access via `inference.allowed_routes`.
193+
- Routes are managed as server-side resources via CLI (`nav inference create/update/delete/list`).
194+
195+
**Inference routes** are stored on the gateway as protobuf objects (`InferenceRoute` in `proto/inference.proto`) and have these fields: `routing_hint` (name for policy matching), `base_url` (backend endpoint), `protocols` (supported API protocols like `openai_chat_completions` or `anthropic_messages`), `api_key`, `model_id`, and `enabled` flag.
196+
197+
**Components involved:**
198+
199+
| Component | Location | Role |
200+
|---|---|---|
201+
| OPA `network_action` rule | `dev-sandbox-policy.rego` | Returns `inspect_for_inference` when no explicit policy match and inference routes exist |
202+
| Proxy interception | `crates/navigator-sandbox/src/proxy.rs` | TLS-terminates intercepted connections, parses HTTP, calls gateway |
203+
| Inference pattern detection | `crates/navigator-sandbox/src/l7/inference.rs` | Matches HTTP method + path against known inference API patterns |
204+
| gRPC forwarding | `crates/navigator-sandbox/src/grpc_client.rs` | Sends `ProxyInferenceRequest` to the gateway |
205+
| Gateway inference service | `crates/navigator-server/src/inference.rs` | Resolves routes from policy, delegates to router |
206+
| Inference router | `crates/navigator-router/src/lib.rs` | Selects a compatible route by protocol and proxies to the backend |
207+
| Proto definitions | `proto/inference.proto` | `InferenceRouteSpec`, `ProxyInferenceRequest/Response`, CRUD RPCs |
208+
209+
161210
### Container and Build System
162211

163212
The platform produces three container images:
@@ -179,7 +228,7 @@ Sandbox behavior is governed by policies written in YAML and evaluated by an emb
179228
- **Filesystem access**: Which directories are readable, which are writable.
180229
- **Network access**: Which remote hosts each program in the sandbox can connect to, with per-binary granularity.
181230
- **Process privileges**: What user/group the agent runs as.
182-
- **Inference routing**: Which AI model endpoints the agent can access.
231+
- **Inference routing**: Which AI model backends the sandbox can route inference traffic to, referenced by `routing_hint` name.
183232
- **L7 inspection rules**: Protocol-level constraints on HTTP API calls for specific endpoints.
184233

185234
Policies are not intended to be hand-edited by end users in normal operation. They are associated with sandboxes at creation time and fetched by the sandbox supervisor at startup via gRPC. For development and testing, policies can also be loaded from local files.
@@ -242,3 +291,5 @@ This opens an interactive SSH session into the sandbox, with all provider creden
242291
| [Sandbox Connect](sandbox-connect.md) | SSH tunneling into sandboxes through the gateway. |
243292
| [Providers](sandbox-providers.md) | External credential management, auto-discovery, and runtime injection. |
244293
| [Policy Language](security-policy.md) | The YAML/Rego policy system that governs sandbox behavior. |
294+
| [Inference Routing](inference-routing.md) | Transparent interception and rerouting of AI inference API calls from sandboxed agents to policy-controlled backends. |
295+
| [Local Inference Routing Demo](inference-routing-local-demo.md) | Step-by-step recording script for showing OpenAI SDK interception and reroute to a local LM Studio backend. |

0 commit comments

Comments
 (0)