Skip to content

perf(proxy): lighter request path: alias the message snapshot, keep the token cache on re-counts, decompress off the event loop - #3909

Merged
chopratejas merged 2 commits into
headroomlabs-ai:mainfrom
bill3wits:perf/o1-o2-o3-snapshot-alias-count-cache-offload
Oct 4, 2026
Merged

chopratejas merged 2 commits into
headroomlabs-ai:mainfrom
bill3wits:perf/o1-o2-o3-snapshot-alias-count-cache-offload

Conversation

@bill3wits

Copy link
Copy Markdown
Contributor

Description

Hi! Three small costs on every proxied request, found while running headroom in front of a busy agent loop:

  • Snapshot copy. Both handlers copy.deepcopy the client messages before the pipeline, which already makes its own deep copy, so long conversations were copied twice per request. The snapshot now aliases the list (HEADROOM_SNAPSHOT_DEEPCOPY=1 brings the copy back).
  • Token cache wiped on re-counts. TokenCountCache.put ran the admission check even for a key already cached, so a full cache cleared itself when an agent re-counted a stable prefix. A put on a known key now just updates it; a new key past the cap still clears, as before.
  • Decompression on the event loop. zstd/gzip/deflate/br bodies were decompressed inline. They now go through one bounded helper on asyncio.to_thread; the [security] Unbounded decompression of request bodies enables zip-bomb denial of service #3284 size caps and the 413/400 errors are unchanged.

Two small things ride along: both shared body readers strip streaming-only index keys from content blocks (well-formed requests pass through byte-identical), and a dead except (ruff B025) that let a missing codec escape as a raw ImportError instead of a 400.

Happy to split this. It is three independent changes plus two small fixes, and CONTRIBUTING asks for one logical change per PR. If you'd rather review them separately, say which and I'll open them one at a time. Opening as a draft until you've had a look.

Type of Change

  • Performance improvement
  • Bug fix (non-breaking change that fixes an issue): the B025 codec path and the index strip

Changes Made

  • proxy/handlers/anthropic.py, openai.py: snapshot alias + env switch
  • tokenizers/base.py: early return in TokenCountCache.put for a cached key
  • proxy/helpers.py: _decompress_bounded on a worker thread; readers call the index strip
  • utils.py: strip_streaming_only_content_fields_in_place
  • tests: 8 new (test_output_only_request_blocks.py, test_token_count_cache.py)

Testing

  • Unit tests pass (pytest): the proxy suite and the changed-area files, not the whole suite
  • Linting passes (ruff check, ruff format --check) on the touched files
  • Type checking (mypy headroom): not run
  • New tests added

Test Output

# Linux, Python 3.12.13, source checkout + the compiled core from the published wheel (not built from source)
$ pytest -q tests/test_proxy*.py tests/test_openai_message_level_cache_control.py tests/test_request_body_decompression_limits.py \
    tests/test_output_only_request_blocks.py tests/test_token_count_cache.py tests/test_utils.py tests/test_tokenizer.py tests/test_tokenizers.py
this branch: 1095 passed, 123 skipped, 2 failed
main:        1087 passed, 123 skipped, 2 failed
# the same two fail on main in this environment (probably the prebuilt core vs. source), untouched by this PR:
#   tests/test_proxy_openai_cache_stability.py::test_openai_chat_outcome_provider_uses_the_fixed_taxonomy
#   tests/test_proxy_openai_cache_stability.py::test_openai_chat_custom_base_flood_cannot_grow_the_provider_set
$ ruff check <7 files>            All checks passed!
$ ruff format --check <7 files>   7 files already formatted

Real Behavior Proof

  • Environment: Linux, Python 3.12.13. headroom proxy --host 127.0.0.1 --port 18787 --no-telemetry run from this branch, ANTHROPIC_TARGET_API_URL pointed at a small local stand-in for /v1/messages that records what the proxy forwards. Same run repeated on main for comparison.
  • Exact steps: (1) after an 8 s warm-up, POST /v1/messages with a gzip-encoded body (228 kB on the wire, 77.4 MB decompressed, 79 messages) while polling /livez every 10 ms; (2) POST a message whose text block carries a stray "index": 0.
  • Observed result (this branch):
    gzip body: 228 kB on the wire, 77.4 MB decompressed -> HTTP 200 in 1.29s
    upstream received: 77,422,858 bytes, 79 messages
    /livez while idle:          355 polls, max 2.2 ms, over 50 ms: 0
    /livez during that request: 28 polls, max 288.3 ms, over 50 ms: 5, p50 2.3 ms
    content block with a stray `index`: HTTP 200; upstream saw no `index` key
    
    Main, same steps: HTTP 200 in 1.35s; /livez during the request max 303.3 ms, 6 polls over 50 ms; no index forwarded either (the Anthropic handler already strips it on that path).
  • Honest reading: the decompression path works end to end off the loop, with the caps intact. At this size the loop stall is dominated by JSON parsing and the pipeline, not decompression, so this run shows no measurable latency difference. I'm not claiming a latency win from it. The snapshot alias and the cache change remove work, but they weren't benchmarked here.
  • Not tested: a real provider under production load; mypy; the full suite; the OpenAI and passthrough paths live (unit tests only).

Runtime Rollout Safety

  • Stable/default behavior changed: snapshot aliases by default; cache keeps entries on re-count; decompression on a worker thread; index stripped at parse
  • Kill switch: HEADROOM_SNAPSHOT_DEEPCOPY=1 for the snapshot; the rest have none
  • Rollback path: revert the PR

Review Readiness

  • I have performed a self-review
  • This PR is ready for human review (draft: waiting on whether you'd like it split)

Additional Notes

Worth a reviewer's eye: with the alias, the snapshot shares the live messages list. Nothing between the snapshot and its readers mutates that list in place today, and a code comment warns against starting to.

@github-actions

github-actions Bot commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

PR governance

This PR does not yet satisfy the required template fields:

  • Fill in Real Behavior Proof → Environment.
  • Fill in Real Behavior Proof → Exact command / steps.
  • Fill in Real Behavior Proof → Observed result.
  • Fill in Real Behavior Proof → Not tested.
  • Fill in Runtime Rollout Safety → Rollout-managed feature(s).
  • Fill in Runtime Rollout Safety → Minimum rollout channel.
  • Fill in Runtime Rollout Safety → Kill switch / disable path.
  • Fill in Runtime Rollout Safety → Unsafe override required.
  • Fill in Runtime Rollout Safety → Qualification impact.
  • Check This PR is ready for human review or convert the PR back to draft.

Please update the PR body, or move the PR back to draft while it is still in progress.

@github-actions github-actions Bot added status: needs author action Pull request body or readiness checklist still needs author updates status: needs rebase Pull request branch is behind the base branch on files it also changes and removed status: needs rebase Pull request branch is behind the base branch on files it also changes labels Oct 1, 2026

@JerrettDavis JerrettDavis left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed refreshed head fff18bd after updating from main. The token-cache and bounded decompression changes have useful focused coverage, but the snapshot alias needs an ownership fix before approval.

original_client_messages = messages is assigned before supported config.hooks.pre_compress calls. The hook contract explicitly permits modifying the messages list, and both Anthropic token mode and OpenAI chat pass it directly to the hook. A hook that edits messages[i]["content"] in place now edits the original snapshot too. That snapshot is subsequently used for session/lineage and replay bookkeeping, so it no longer preserves the documented pre-hook client history. The pipeline's own deepcopy happens after the hook and cannot protect this earlier boundary; extension mutation surfaces also need to be considered.

Please retain independent snapshot ownership when mutation-capable hooks/extensions are configured, or supply those stages with a separately owned message graph. Add a handler regression with an in-place hook that verifies the recorded original messages remain the pre-hook input, while the forwarded messages contain the hook result. Cover OpenAI's default-cache hook path as well as Anthropic token mode. An optional environment escape hatch does not preserve the existing default contract.

Validation: 28 selected ingress/token-cache/body-size tests passed with the existing native extension. No contributor implementation was modified.

@JerrettDavis JerrettDavis left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Re-reviewed current head 32ec1a9, which includes main. The prior snapshot-ownership blocker is resolved: configured hooks and enabled extensions now receive an independently owned original snapshot, with regressions confirming pre-hook originals and post-hook forwarded content for OpenAI and Anthropic. All 29 focused snapshot/ingress/token-cache tests passed with the native extension; pinned Ruff check/format and diff checks passed.

I authorized the five pending hosted workflows for this head. Commitlint now reports a concrete required-gate failure: the original commit subject perf(proxy): alias the client-message snapshot, stop re-counts clearing the token cache, decompress bodies off the event loop is 125 characters, exceeding the repository's 100-character header limit. Please shorten that subject (or squash into a compliant subject) without changing the code. I will authorize and inspect CI for the resulting head; no contributor CI action is needed. Other workflows are still running, so their final outcome is not yet established.

@codecov-commenter

codecov-commenter commented Oct 3, 2026 •

Copy link
Copy Markdown

⚠️ Please install the 'codecov app svg image' to ensure uploads and comments are reliably processed by Codecov.

Codecov Report

❌ Patch coverage is 59.15493% with 29 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
headroom/proxy/helpers.py 42.55% 26 Missing and 1 partial ⚠️
headroom/utils.py 84.61% 0 Missing and 2 partials ⚠️

📢 Thoughts on this report? Let us know!

… loop

O1: the pre-pipeline snapshot of the client messages aliases the live list
instead of copy.deepcopy when nothing can mutate it; the pipeline already
works on its own deep copy. snapshot_original_messages() returns an
independently owned deep copy whenever config.hooks is set or a pipeline
extension is enabled (pre_compress receives the live list and may edit it in
place, and the recorded original feeds session/lineage and replay
bookkeeping), and keeps the alias only when neither is configured. Covers
the Anthropic messages and batch paths and the OpenAI chat path, including
the re-snapshot after an INPUT_RECEIVED extension replaces the list. No env
switch.

O2: TokenCountCache.put on a text that is already cached updates the value
and returns before the admission check. A full cache used to clear on that
put, so re-counting a known prefix forced a full re-encode on the next miss.
Admitting a new key past the cap still clears, as before.

O3: Content-Encoding request bodies (zstd, gzip, deflate, br) are
decompressed in asyncio.to_thread through one dispatch point,
_decompress_bounded. RequestBodyTooLarge and the "Failed to decompress ..."
errors propagate unchanged (413 / 400 as before). A missing optional codec
is a 400 again: the first except clause caught (ValueError, ImportError) and
re-raised, so the ImportError clause below it could never run (ruff B025);
it now catches ValueError only.

Also: both shared body readers strip streaming-only `index` keys from
content blocks in place (strip_streaming_only_content_fields_in_place).
Well-formed requests pass through byte-identical; malformed ones that echo
`index` no longer draw an upstream 400. The Anthropic handler's own strip
stays (idempotent) for paths that bypass the readers.

Tests: index strip and token-cache re-count paths (8 new); handler
regressions with an in-place pre_compress hook for OpenAI's default-cache
path and Anthropic token mode (the original stays pre-hook at
compute_session_id, the forwarded body carries the hook result); the
alias/copy rule. Ruff format on the touched files.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@bill3wits
bill3wits force-pushed the perf/o1-o2-o3-snapshot-alias-count-cache-offload branch from 32ec1a9 to f75e854 Compare October 3, 2026 14:52

@JerrettDavis JerrettDavis left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Rechecked f75e854 after the commit rewrite. It already contains current main, and its tree is identical to the previously reviewed 32ec1a9 head. The new 71-character subject resolves my remaining commitlint blocker. The earlier ownership fix remains intact; no further introduced correctness blocker found in the unchanged code.

Exact-head hosted checks are terminal with no failures. No merge has been performed.

@bill3wits
bill3wits marked this pull request as ready for review October 3, 2026 18:03

@devin-ai-integration devin-ai-integration Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Devin Review found 4 potential issues.

Devin Review

Comment thread headroom/proxy/helpers.py
# is unchanged.
from headroom.utils import strip_streaming_only_content_fields_in_place

strip_streaming_only_content_fields_in_place(result.get("messages"))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Bedrock still forwards invalid index keys

When a Bedrock content block carries index, read_request_json_with_bytes removes it only from parsed messages. The Bedrock forwarder sends unchanged raw bytes on bypass, so the upstream still receives the rejected field.

Learn more

The bytes-returning reader returns both parsed messages and original bytes. The new canonicalizer modifies only the parsed messages, but handle_bedrock_invoke starts with the returned bytes as its outbound body and keeps them on bypass or unchanged compression. A client replaying a streaming response block with index thus sends the same invalid key upstream despite the new normalization.

Example: A Bedrock invoke request with messages: [{"role":"assistant","content":[{"type":"text","text":"hi","index":0}]}] and x-headroom-bypass: true yields parsed content without index, but forwards bytes containing "index":0.

Recommended fix: Have strip_streaming_only_content_fields_in_place signal whether it changed a message. In read_request_json_with_bytes, re-encode returned bytes whenever stripping occurred, as already done for output-only blocks, while preserving unchanged requests byte-for-byte. Check other consumers of the returned raw bytes.

Devin Review


Was this helpful? React with 👍 or 👎 to provide feedback.

Comment thread headroom/proxy/helpers.py
# requests.
from headroom.utils import strip_streaming_only_content_fields_in_place

strip_streaming_only_content_fields_in_place(result.get("messages"))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Batch items retain invalid index keys

For Anthropic batches, strip_streaming_only_content_fields_in_place sees no top-level messages; items live under requests[].params.messages. The batch handler forwards those indices unchanged, so affected batch items fail upstream.

Learn more

The shared bytes-less reader now normalizes result.get("messages"), but Anthropic's batch endpoint stores each item's conversation in requests[].params.messages. The batch handler does not run its own index canonicalizer. As a result, the newly introduced normalization does not apply to batch items, including echoed streaming content blocks that Anthropic rejects.

Example: A batch containing {"requests":[{"custom_id":"a","params":{"model":"claude-sonnet-4-5","messages":[{"role":"assistant","content":[{"type":"text","text":"hi","index":0}]}]}}]} retains index in the forwarded item.

Recommended fix: Normalize each params.messages inside the Anthropic batch handler before processing and forwarding, including the no-optimization and fail-open branches. Keep changes scoped to message content, not arbitrary batch metadata.

Devin Review


Was this helpful? React with 👍 or 👎 to provide feedback.

Comment thread headroom/proxy/helpers.py
Comment on lines +4692 to +4694
if hooks is not None or bool(getattr(extensions, "enabled", False)):
return copy.deepcopy(messages)
return messages

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔍 Advertised snapshot kill switch is absent

The rollout guidance advertises HEADROOM_SNAPSHOT_DEEPCOPY=1, but snapshot_original_messages never reads it. Confirm the intended rollback control and align the code or guidance.

Devin Review


Was this helpful? React with 👍 or 👎 to provide feedback.

Comment thread headroom/proxy/helpers.py
Comment on lines +3157 to +3161
raw = await asyncio.to_thread(
_decompress_bounded,
raw,
encoding,
)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔍 Independent changes bundled into one PR

Snapshot aliasing, token-cache admission, decompression offloading, and content normalization are independent changes. CONTRIBUTING.md requires one logical change per PR; consider splitting them.

Devin Review


Was this helpful? React with 👍 or 👎 to provide feedback.

@chopratejas
chopratejas merged commit c873d13 into headroomlabs-ai:main Oct 4, 2026
41 checks passed
JerrettDavis added a commit that referenced this pull request Oct 6, 2026
🤖 I have created a release *beep* *boop*
---


##
[0.40.0](v0.39.1...v0.40.0)
(2026-10-06)


### ⚠ BREAKING CHANGES

* **proxy:** read HEADROOM_LICENSE and make usage reporting opt-in
([#3857](#3857))
* drop the crewai extra to remove chromadb from the lockfile
([#3870](#3870))

### Features

* **compress:** add per-message compression diagnostics
([#3058](#3058))
([fef99cc](fef99cc))
* **compress:** live-agent densify mode, lossless and cache-safe
([#1402](#1402))
([84f0e84](84f0e84))
* **dashboard:** add CO₂ Saved card to dashboard and /stats API
([#1369](#1369))
([2bc9419](2bc9419))
* **dashboard:** give savings metrics one canonical home
([#3320](#3320))
([6af7efe](6af7efe))
* **learn:** add agy (Antigravity CLI) as an analysis backend
([#3939](#3939))
([4ba231a](4ba231a))
* **live-zone:** wire the SourceCode and PlainText dispatch arms
([#3227](#3227))
([49f69be](49f69be))
* **memory:** persist TrafficLearner pending evidence across restarts +
expose learner stats
([#3104](#3104))
([4257ed4](4257ed4))
* **proxy:** add x-headroom-keep-last-turns header for per-request
context trimming (issue
[#2858](#2858))
([#3059](#3059))
([246162d](246162d))
* **proxy:** route Gemini plugin traffic through native transforms
([#2697](#2697))
([0ad996a](0ad996a))
* **proxy:** serve on a Unix domain socket (--uds)
([#3151](#3151))
([e7b3baf](e7b3baf))
* **wrap:** add headroom wrap bob for IBM Bob CLI
([#3801](#3801))
([0b7dab4](0b7dab4))
* **wrap:** extend reduce-at-source quiet defaults to telemetry/nag
banners
([#2550](#2550))
([3a3a465](3a3a465))


### Bug Fixes

* **anthropic:** preserve failed CCR continuations
([#3843](#3843))
([94bb055](94bb055))
* **backends/anyllm:** map Anthropic tool_choice none to OpenAI none
([#3963](#3963))
([d722a64](d722a64))
* **backends/litellm:** align streaming usage tokens with the
non-streaming path
([#2688](#2688))
([90a68e0](90a68e0))
* **backends/litellm:** keep the upstream 4xx status on send_message
errors
([#3944](#3944))
([e0e41cd](e0e41cd))
* **backends/litellm:** map Anthropic tool_choice none to OpenAI none
([#2689](#2689))
([57a5708](57a5708))
* **backends/litellm:** match the caller-key Bearer scheme
case-insensitively (RFC 7235)
([#3965](#3965))
([848d7a0](848d7a0))
* **backends/litellm:** report requested model in OpenAI streaming
chunks
([#2690](#2690))
([8538a83](8538a83))
* **cache:** keep prefix lineage across a replaced system tail
([#3934](#3934))
([16eaec1](16eaec1))
* **ccr:** answer headroom_retrieve on the direct chat path instead of
forwarding it
([#3816](#3816))
([126e144](126e144))
* **ccr:** inject headroom_retrieve before the prefix is warm, not after
([#3810](#3810))
([bb8c285](bb8c285))
* **ccr:** only proactively expand compressions present in the
requesting conversation
([#3924](#3924))
([9b9a883](9b9a883))
* **ccr:** unwrap Hermes batched tool_call so headroom_retrieve stays
exempt
([#3839](#3839))
([2f07668](2f07668))
* **ci:** clear dependency audit and gate Docker publishing
([#3984](#3984))
([67ce7d0](67ce7d0))
* **ci:** keep hard watchdog out of pytest shards
([#3845](#3845))
([2527431](2527431))
* **ci:** tolerate missing Docker cache blobs
([#3945](#3945))
([540f4da](540f4da))
* clarify CCR marker content preservation
([#3175](#3175))
([005a4e1](005a4e1))
* **cli:** honor CODEX_HOME and align wrap/init/doctor with the live
proxy port
([#3855](#3855))
([5bf6612](5bf6612))
* **cli:** strip unmarked headroom_memory TOML table before injecting
([#3490](#3490))
([08dda4c](08dda4c))
* **codex:** explain why wrapped Codex runs without the shared server
([#3899](#3899))
([1982b3a](1982b3a))
* **codex:** preserve remote compaction in init config
([#3410](#3410))
([6df4a96](6df4a96))
* **codex:** resolve per-turn project context
([#2636](#2636))
([0396ea2](0396ea2))
* **codex:** strip stale compression framing on the Responses subpath
passthrough
([#3792](#3792))
([0eba65c](0eba65c))
* **copilot:** defer keychain auth lookup
([#2739](#2739))
([81a8a28](81a8a28))
* **copilot:** read timezone-naive token expiry as UTC, not host local
time ([#3216](#3216))
([2b2dc1b](2b2dc1b))
* **copilot:** warn when a VS Code profile cannot see the proxy settings
([#3919](#3919))
([7b68fc2](7b68fc2))
* **deps:** patch brace-expansion in wrap E2E lockfile
([#3873](#3873))
([5e0435d](5e0435d))
* **deps:** patch urllib3 and Next.js advisories
([#3896](#3896))
([573e385](573e385))
* **doctor:** recognize Azure Foundry routing for Claude Code
([#1339](#1339))
([d5318ac](d5318ac))
* **grok:** route CLI traffic to api.x.ai on shared proxies
([#2693](#2693))
([ef1c528](ef1c528))
* honor configured port in container startup and healthcheck
([#2436](#2436))
([1b7977c](1b7977c))
* **install/apply:** add --no-rate-limit flag, persist in proxy_args
([#1350](#1350))
([#1365](#1365))
([31d0344](31d0344))
* **install:** parse Windows Task Scheduler XML output
([#3830](#3830))
([c32a4f4](c32a4f4))
* **install:** replace the launchd wrapper with the proxy listener
([#3224](#3224))
([44db755](44db755))
* **install:** use safe model backends for persistent services
([#3646](#3646))
([7287589](7287589))
* **integrations:** record metric timestamps in UTC, not naive local
time ([#3117](#3117))
([76ef2c3](76ef2c3))
* **kompress:** degrade when a native dep is installed but unloadable
([#3133](#3133))
([8f3d677](8f3d677))
* **kompress:** pin ModernBERT tokenizer and encoder revisions
([#3808](#3808))
([2de5828](2de5828))
* **learn:** drop the echoed prompt from a failed cli's error
([#3928](#3928))
([f00d425](f00d425))
* **learn:** keep CLAUDE.local.md out of git via .git/info/exclude
([#3108](#3108))
([c072251](c072251))
* **learn:** keep transcript-derived content inside the managed block
([#3850](#3850))
([c46e74d](c46e74d))
* **learn:** merge git worktree sessions into the repo's project
([#3854](#3854))
([3fdf18e](3fdf18e))
* **learn:** preserve stable traffic pattern items
([#2293](#2293))
([1631cde](1631cde))
* **learn:** recognize escaped persisted pattern IDs
([#3951](#3951))
([2e74f06](2e74f06))
* **learn:** run the claude-cli analysis with hooks disabled
([#3926](#3926))
([2a4d34c](2a4d34c))
* **learn:** run the claude-cli analysis with no tools
([#3892](#3892))
([db90b93](db90b93))
* **log_compressor:** keep and name pytest short-summary failures
([#3828](#3828))
([d90b320](d90b320))
* make Headroom work behind Zscaler and other TLS-inspecting networks
([#3831](#3831))
([66258c4](66258c4))
* **mcp:** bound version-detection git subprocess
([#3038](#3038))
([0712e04](0712e04))
* **mcp:** escape control characters when rendering TOML server blocks
([#3964](#3964))
([a96154f](a96154f))
* **mcp:** keep other apps' tables when replacing the Codex/Grok MCP
span ([#3877](#3877))
([d0fd56e](d0fd56e))
* **mcp:** resolve opencode.jsonc for MCP registration
([#2496](#2496))
([f824a27](f824a27))
* **memory:** align project routing with wrap headers
([#3603](#3603))
([b1b005a](b1b005a))
* **memory:** avoid injecting tools into tool-free requests
([#3677](#3677))
([46755b5](46755b5))
* **memory:** close SQLite connections in memory/fts5/graph adapters
([#3153](#3153))
([231a627](231a627))
* **memory:** handle list-shaped system content in inline memory
injection
([#3794](#3794))
([717527b](717527b))
* **memory:** ignore leading cd prefix when pairing Bash error
recoveries
([#3776](#3776))
([143a38d](143a38d))
* **memory:** pin LF on memory writers and guard the learn-writer
newline contract
([#3706](#3706))
([b10dd8d](b10dd8d))
* **oauth2:** mint and inject only on requests that go upstream
([#3849](#3849))
([f977d52](f977d52))
* **offline:** make HEADROOM_OFFLINE a real air-gap via one chokepoint
([#3729](#3729))
([119d1a1](119d1a1))
* **opencode:** hide spawned Windows console windows
([#3743](#3743))
([2157400](2157400))
* **opencode:** route only LLM traffic through Headroom
([#3884](#3884))
([d75eecd](d75eecd))
* **output-savings:** seed the holdout key on the whole first user
message
([#3209](#3209))
([117ff72](117ff72))
* **parser:** stop counting HTML comments twice in waste signals
([#3942](#3942))
([d0e9c4c](d0e9c4c))
* **parser:** whitespace waste signal always reported zero
([#1102](#1102))
([f19bc9a](f19bc9a))
* **plugin:** resolve hook CLI through plugin-root launcher
([#3053](#3053))
([ef7605f](ef7605f))
* **plugins/openclaw:** return messages the proxy did not change exactly
as they came in
([#3826](#3826))
([d1ad189](d1ad189))
* prevent HF tokenizer downloads in offline mode
([#3783](#3783))
([d503c57](d503c57))
* **pricing:** add Claude Sonnet 5.5 / Opus 5.5 / Fable 5.1, correct
Sonnet 5 rates
([#3841](#3841))
([0d99c56](0d99c56))
* protect file reads in chained shell commands
([#2668](#2668))
([ffc3599](ffc3599))
* **providers/vertex:** Vertex route multi-region locations
([#3802](#3802))
([a9c1ac5](a9c1ac5))
* **proxy/anthropic:** keep tool_result blocks first when neutralizing
headroom_retrieve history
([#3874](#3874))
([ccd9fff](ccd9fff))
* **proxy/batch:** honor x-headroom-bypass on the batch paths
([#2570](#2570))
([9b26a49](9b26a49))
* **proxy:** add same-origin check to /v1/retrieve/tool_call
([#3955](#3955))
([2297b50](2297b50))
* **proxy:** bill the whole prompt on the gateway path so savings read
true ([#3632](#3632))
([befdb52](befdb52))
* **proxy:** carry safeguard_results through SSE resynthesis
([#3958](#3958))
([f37ef59](f37ef59))
* **proxy:** classify Pi Codex Responses alias
([#2583](#2583))
([a493f55](a493f55))
* **proxy:** do not cache error replies delivered as http 200
([#3930](#3930))
([1132a54](1132a54))
* **proxy:** drop a tool_reference naming the search tool itself
([#3172](#3172))
([b73adaa](b73adaa))
* **proxy:** F3 — per-tenant TOIN learning key
([#404](#404))
([f90a56b](f90a56b))
* **proxy:** forward operator-listed guarded upstreams through a proxy
([#3804](#3804))
([4b7e5d2](4b7e5d2))
* **proxy:** freeze and replay the forwarded prefix on Gemini paths
([#3394](#3394))
([#3865](#3865))
([dd84321](dd84321))
* **proxy:** gate the raw request body, not just its Content-Length
header
([#3338](#3338))
([c946b6b](c946b6b))
* **proxy:** hard watchdog that dumps and exits when a native call
seizes the GIL
([#3180](#3180))
([79daeb5](79daeb5))
* **proxy:** inject memory context past a trailing system message
([#3948](#3948))
([3948dbe](3948dbe))
* **proxy:** keep cache_control outside content blocks where the Chat
Completions client put it
([#3895](#3895))
([a6d6c14](a6d6c14))
* **proxy:** keep exception text and the proxy token out of client
output
([#3851](#3851))
([861e94d](861e94d))
* **proxy:** keep the newest user message verbatim in cache-mode delta
compression
([#3923](#3923))
([613ae92](613ae92))
* **proxy:** key rate limits by peer unless the caller authenticated
([#3860](#3860))
([6fb7cad](6fb7cad))
* **proxy:** kill the image worker a timed-out call abandons
([#3940](#3940))
([793bb85](793bb85))
* **proxy:** label plain OpenAI chat traffic openai, not custom
([#3912](#3912))
([7df8bd8](7df8bd8))
* **proxy:** log response_content_length in proxy_inbound_response
([#2701](#2701))
([cfa479a](cfa479a))
* **proxy:** preserve Bedrock body-limit error dialect
([#3871](#3871))
([ffc6edb](ffc6edb))
* **proxy:** preserve Windows service when Rust core is blocked
([#2989](#2989))
([eaa16d9](eaa16d9))
* **proxy:** quarantine compression only once half the pool is stuck
([#3932](#3932))
([be2b205](be2b205))
* **proxy:** read HEADROOM_LICENSE and make usage reporting opt-in
([#3857](#3857))
([7460389](7460389))
* **proxy:** reap wrap-spawned proxies once no wrap clients remain
([#3202](#3202))
([4227bd2](4227bd2))
* **proxy:** redact upstream error detail and add opt-in /metrics
loopback gate
([#2589](#2589))
([a05717f](a05717f))
* **proxy:** refuse token-less non-loopback binds, gate operator routes
([#3852](#3852))
([ee731d7](ee731d7))
* **proxy:** reset the cc-switch upstream when Claude Official is
selected
([#3166](#3166))
([fed7281](fed7281))
* **proxy:** run extension middleware inside the security gate and body
ceiling
([#3847](#3847))
([1854fd7](1854fd7))
* **proxy:** run injected memory tools server-side on streaming turns
([#3947](#3947))
([1cf4966](1cf4966))
* **proxy:** share one compression deadline across a Responses request
([#3938](#3938))
([4d27d02](4d27d02))
* **proxy:** skip pricing lookup for passthrough:* models
([#2585](#2585))
([f87848c](f87848c))
* **proxy:** stop headroom logger from suppressing propagation to
stdout/stderr
([#3096](#3096))
([f78e66f](f78e66f))
* **relevance:** bound segment size by max_chars
([#2518](#2518))
([0ef7b2a](0ef7b2a))
* **reporting:** price the savings tile instead of fabricating $0.00
([#3821](#3821))
([f519fa8](f519fa8))
* **router:** fall back to built-ins when an external compressor passes
through
([#3893](#3893))
([2c4b20f](2c4b20f))
* **router:** keep ccr_retrieve exemption through orchestrator wrappers
([#3915](#3915))
([6151ed1](6151ed1))
* **rust-proxy:** forward request paths verbatim and pin the rustls
provider
([#3853](#3853))
([9d98ea5](9d98ea5))
* **savings:** record tool-schema dollars disjointly beside the folded
headline
([#3170](#3170))
([c719d4a](c719d4a))
* **savings:** stop scoring unobserved strata against the global mean
([#3128](#3128))
([f864525](f864525))
* **sdk:** keep tool names on Vercel tool-result parts through the
OpenAI round trip
([#3883](#3883))
([b715671](b715671))
* **sdk:** preserve Gemini media parts in message conversion instead of
dropping the turn
([#3882](#3882))
([038c923](038c923))
* **sdk:** stop JSON-wrapping structured tool_result content in the
Anthropic adapter
([#3797](#3797))
([2e4a60a](2e4a60a))
* **search:** stop a context line's body from becoming its line marker
([#3788](#3788))
([f3f2e00](f3f2e00))
* **security:** create memory stores and other state files owner-only
([#3848](#3848))
([5119b6e](5119b6e))
* **security:** exempt only GET health probes from the proxy token
([#3921](#3921))
([d99779d](d99779d))
* **security:** partition the response cache by credential and principal
([#3862](#3862))
([ee6cfcc](ee6cfcc))
* **security:** poll usage only with the operator's own credential
([#3863](#3863))
([65e63c0](65e63c0))
* **security:** stop forwarding the proxy token to upstream providers
([#3891](#3891))
([0147cf0](0147cf0))
* **settings:** report live proxy configuration
([#3177](#3177))
([58b1454](58b1454))
* **smart-crusher:** recurse at adaptive array limit
([#3770](#3770))
([7790bde](7790bde))
* **storage:** page JSONL queries after sorting
([#3872](#3872))
([91237ca](91237ca))
* **subscription:** match the Bearer scheme case-insensitively for the
Codex usage poll (RFC 7235)
([#3966](#3966))
([fe88461](fe88461))
* **subscription:** poll with the refreshed credentials-file OAuth token
([#3916](#3916))
([2dc9fc2](2dc9fc2))
* surface Codex responses traffic in dashboard
([#399](#399))
([ecc4967](ecc4967))
* **telemetry:** label custom-base chat upstreams from a fixed provider
set ([#3759](#3759))
([f0ec2bb](f0ec2bb))
* **thinking:** don't read the model's date suffix as its minor version
([#3791](#3791))
([afaaaa8](afaaaa8))
* **transforms/kompress:** preserve line boundaries and tabular output
in Kompress
([#3119](#3119))
([fe2ed2b](fe2ed2b))
* **transforms:** judge a Codex exec envelope read by its output
([#3878](#3878))
([922924e](922924e))
* **transforms:** judge the code_aware Kompress fallback in tokens
([#3881](#3881))
([6326965](6326965))
* **transforms:** keep record-bearing JSON out of Kompress
([#3673](#3673))
([#3880](#3880))
([8dbbd1d](8dbbd1d))
* **wrap:** never reuse a non-Headroom listener on the proxy port
([#3799](#3799))
([3ffa57e](3ffa57e))
* **wrap:** never reuse a proxy with incompatible routing config
([#3201](#3201))
([f69e246](f69e246))
* **wrap:** preserve pre-set ANTHROPIC_BASE_URL as proxy upstream
([#1358](#1358))
([bb1ab6a](bb1ab6a))
* **wrap:** scope Serena's MCP registration to the wrapped project
([#2787](#2787))
([#2992](#2992))
([8cfeb69](8cfeb69))
* **wrap:** set ANTHROPIC_HOST so `wrap goose` actually proxies
Anthropic
([#2619](#2619))
([c07fad0](c07fad0))
* **wrap:** strip -dev from running proxy version in restart check
([#3200](#3200))
([bd0296b](bd0296b))


### Performance Improvements

* add savings audit output
([#1211](#1211))
([9a72bc0](9a72bc0))
* **compression/code:** avoid a UTF-8 copy per code block in
byte-to-char mapping
([#3169](#3169))
([d5e5534](d5e5534))
* **image:** OCR and SigLIP each image once, not once per turn
([#3941](#3941))
([a8561fb](a8561fb))
* **memory:** bound traffic-learner _persisted_ids with the dedup window
([#3341](#3341))
([7e73438](7e73438))
* **metrics:** cap inbound request-path cardinality to bound memory
([#3340](#3340))
([88a1f4e](88a1f4e))
* **proxy:** lighter request path: alias the message snapshot, keep the
token cache on re-counts, decompress off the event loop
([#3909](#3909))
([c873d13](c873d13))
* **tokenizer:** memoize OpenAI token counts like the Anthropic counter
([#3168](#3168))
([0a2c80d](0a2c80d))


### Dependencies

* bump pyjwt 2.13.0 -&gt; 2.15.1 for CVE-2026-102274
([#3861](#3861))
([37b9c46](37b9c46))
* Bump source-map-js from 1.2.1 to 1.2.2 in /docs
([#3986](#3986))
([e06631a](e06631a))
* Bump source-map-js from 1.2.1 to 1.2.2 in /plugins/openclaw
([#3988](#3988))
([4be83b3](4be83b3))
* Bump source-map-js from 1.2.1 to 1.2.2 in /plugins/opencode
([#3987](#3987))
([478dd9e](478dd9e))
* Bump source-map-js from 1.2.1 to 1.2.2 in /sdk/typescript
([#3985](#3985))
([62cde35](62cde35))
* bump the npm-minor-patch group across 2 directories with 11 updates
([#3903](#3903))
([afe4f4e](afe4f4e))
* bump the npm-minor-patch group across 3 directories with 11 updates
([#3910](#3910))
([ff1a0d6](ff1a0d6))
* bump webpki-roots from 0.26.11 to 1.0.8
([#3905](#3905))
([fee53b4](fee53b4))
* drop the crewai extra to remove chromadb from the lockfile
([#3870](#3870))
([59b8cef](59b8cef))


### Code Refactoring

* **ccr:** read the expansion query through extract_user_query
([#2707](#2707))
([dcec845](dcec845))
* **wrap:** generate goose/openhands/openclaude from a WrapTarget
registry
([#3800](#3800))
([de4cf7e](de4cf7e))

---
This PR was generated with [Release
Please](https://github.com/googleapis/release-please). See
[documentation](https://github.com/googleapis/release-please#release-please).

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

status: needs author action Pull request body or readiness checklist still needs author updates

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants