Skip to content

fix: restore implicit-cache hit rate — per-conversation session IDs + persistent signature cache - #367

Open
tracycam wants to merge 5 commits into
badrisnarayanan:mainfrom
tracycam:fix/implicit-cache-session-affinity
Open

tracycam wants to merge 5 commits into
badrisnarayanan:mainfrom
tracycam:fix/implicit-cache-session-affinity

Conversation

@tracycam

@tracycam tracycam commented Sep 3, 2026

Copy link
Copy Markdown

Problem

Two independent bugs pin the Gemini implicit-cache hit rate at ~45-55% for multi-conversation / long-running proxy deployments, even though the backend happily serves 240k+ token prefix hits (e.g. CLIProxyAPI#5226 logs 239k/243k cached):

1. All conversations share one cache-affinity namespace

The Cloud Code backend scopes implicit-cache affinity by session ID, but deriveSessionId() returned one process-lifetime ID per account — every conversation through that account shared it. Interleaved conversations (multiple Claude Code sessions, subagents, parallel clients) each evicted the others' cached KV blocks.

The docblock and request-builder.js comment already described the intended design ("Use stable session ID derived from first user message for cache continuity") — the implementation never did it. Same bug class as router-for-me/CLIProxyAPI#592, where the equivalent fix restored cache hits.

Measured (warmgate telemetry in front of this proxy, single account):

scenario hit rate
one conversation, sequential turns, quiet window ~95% effective (full prefix minus the backend's ~2.5-4k fresh window)
two conversations interleaved (60k + 251k prefixes) 44-55%, pinned — displaced oldest blocks never come back

Byte-level diffs of consecutive requests confirmed the client-side prefix itself was stable; the loss was purely backend eviction.

2. The 2h signature TTL silently rewrites history

Signatures are replayed verbatim on every subsequent turn. After 2h:

  • tool_use parts degrade to the skip_thought_signature_validator marker,
  • thinking blocks get dropped entirely (unknown signature family),

both of which rewrite conversation history mid-session and bust the prefix from that point on. A proxy restart had the same effect since the caches were memory-only.

Changes

session-manager.js — deriveSessionId() now derives a deterministic per-conversation ID from sha256(accountEmail, first user text) (first user text is found even when earlier user messages are tool_result-only), shaped like the binary's uuid + Date.now() format so the backend sees a familiar shape. Stable across turns and proxy restarts. No-user-text requests fall back to the previous per-account behavior.

signature-cache.js — no TTL (signatures live for the process lifetime) plus on-disk persistence: ~/.config/antigravity-proxy/signature-store.json, atomic tmp+rename writes, mode 0600, debounced flush (1s) with periodic (30s) and SIGTERM/SIGINT/exit safety flushes. ANTIGRAVITY_SIGNATURE_STORE overrides the path. Loads at boot so restarts no longer rewrite long-running conversations. All persistence failures fail open — store I/O can never break forwarding.

server.js — /test/clear-signature-cache clears both caches and removes the persisted store.

Cross-model safety is unchanged: Gemini replays are still gated by the signature-family lookup; only the artificial expiry is gone.

Testing

  • tests/test-session-manager.cjs — 9 unit tests: stability, conversation/account separation, tool_result-preamble handling, ID shape, cross-process determinism, fallbacks.
  • tests/test-signature-cache.cjs — 7 unit tests: round trips, no-expiry, 0600 store file, cross-process reload (restart simulation), clear-all removes the store.
  • Existing suites pass (test-strategies.cjs 89/89, test-cache-control.cjs all families).
  • Live verification against the real backend: after the session-ID fix, two interleaved ~60k-token conversations both converged to full-prefix hits (~57.3k cached each turn, only the fresh window uncached); previously the second conversation lost its oldest blocks permanently.

…versation cache eviction

The Cloud Code backend scopes implicit-cache affinity by session ID.
deriveSessionId() returned one process-lifetime ID per account, so every
conversation routed through the same account shared a single cache
namespace. When multiple conversations interleave (multiple Claude Code
sessions, subagents, parallel clients), each large prefix evicts the
others' cached KV blocks.

Measured on a real deployment: sequential turns in a quiet window hit
the full prefix (~95% effective hit rate, backend fresh window only),
while two interleaved conversations dropped to 44-55% and stayed pinned
there as the displaced oldest blocks never returned. The docblock and
request-builder comment already described per-conversation derivation
("derived from the first user message for cache continuity") but the
implementation never did it — same class of bug as CLIProxyAPI#592.

deriveSessionId() now hashes (accountEmail, first user text) into a
deterministic ID shaped like the binary's uuid+Date.now() format, so
each conversation keeps its own affinity namespace across turns and
proxy restarts. Requests with no user text at all fall back to the
previous per-account behavior.
Signatures are replayed verbatim on every subsequent turn of a
conversation. The 2h TTL silently degraded them: tool_use parts fell
back to the skip marker and thinking blocks were dropped entirely once
their family entry expired, rewriting conversation history mid-session
and busting the upstream implicit-cache prefix from that point on. A
proxy restart had the same effect because the caches were in-memory
only.

Signatures are now kept for the whole process lifetime and persisted to
~/.config/antigravity-proxy/signature-store.json (atomic tmp+rename,
mode 0600, debounced flush plus periodic/exit safety flushes, override
with ANTIGRAVITY_SIGNATURE_STORE). The store loads at boot, so restarts
no longer rewrite long-running conversations. All persistence failures
fail open — store I/O can never break request forwarding.

/test/clear-signature-cache now clears both caches and removes the
persisted store. Cross-model safety is unchanged: the family lookup
still gates Gemini replays, only the artificial expiry is gone.
Google streams usageMetadata progressively: early chunks carry an
interim promptTokenCount and no cachedContentTokenCount yet, so
message_start latches input_tokens = (nearly-full prompt) - 0 whenever
the request is served mostly from cache. message_delta then corrected
cache_read_input_tokens but never input_tokens, leaving clients that
combine the two (e.g. gateways computing total = input + cache_read)
with a phantom ~2x prompt and a fabricated ~50% cache-hit rate.

Replay-proof against a captured 136k-token conversation: same body,
stream:true reported input=134,572 + cache=130,467 (phantom 265k total,
'49% hit') while stream:false reported the true 136,389 total with
130,467 cached (95.8% hit). The final usageMetadata chunk is
authoritative, so message_delta now also carries
input_tokens = promptTokenCount - cachedContentTokenCount (same formula
the non-streaming path already uses; Math.max guards odd backends).
Verified end-to-end after the fix: warm repeats report ~5.7k input +
~126k cached (95.7% hit) and the next turn ~5.4k + ~130k (96.0%).
…tch (Gemini only)

The synthetic-message injection ('[N tool executions completed.]' +
'[Continue]', and the mid-conversation splice for interrupted tools)
was added wholesale in 426acc4 with no issue reference and no recorded
upstream error. A/B replay against the live Cloud Code backend shows
Google accepts raw mid-tool-loop, parallel-tool, and interrupted-tool
histories natively and answers identically without the injection, so
for Gemini targets the fabrication is pure downside: a fake user
'[Continue]' becomes the model's most recent instruction on nearly
every adaptive-thinking agentic request, and the toolResultCount-bearing
text differs between requests.

Claude targets keep the recovery because the Anthropic API
hard-rejects conversations with unclosed tool_use ids, so the
interrupted-tool synthesis may still be required there.
Anthropic semantics bill all generated tokens (thinking included) as
output_tokens; the backend reports internal reasoning separately as
thoughtsTokenCount and bills it as output quota. Mapping only
candidatesTokenCount silently undercounts thinking-heavy responses by
40-60% (measured: determinant task reported 1411 while generating ~2400
tokens). Fix both the streaming and non-streaming usage paths.
ameerchik6 added a commit to ameerchik6/antigravity-claude-proxy that referenced this pull request Oct 4, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant