Skip to content

Eliminate first-turn warming retries and preserve the terminal transcript #22552

Description

@standujar

GOAL F — A first chat turn should not cost a retry cycle

Roughly 30 % of shared-chat streaming sends fail and are retried by the client. Users rarely see an
error; they pay for it in time to first reply.

Parent issue: #22508 · scope: shared runtime + Durable Object.

Never declare this goal done on the basis of

  • the client's retry succeeding — that is what hides this today;
  • pointing at shared-runtime-conversation.ts:1764 and writing "cache warming". True, useless, and
    it leaves the problem untouched;
  • absorbing the failure server-side — retrying the DO fetch inside coordinateSharedStream, or
    awaiting pending hydration instead of throwing. The client-visible rate drops and the user still
    waits;
  • a re-measurement that is within noise. n=61 of 199 gives a 95 % CI of roughly ±6.4 points on a
    30 % rate, so a result of 24 % proves nothing. State a target and a sample size before measuring;
  • Worker-side and DO-side counts taken from the observability API and compared to each other — that
    API drops events class-dependently, so the comparison is not sound in either direction.

Exact result to prove

  1. The failing turn no longer fails — not the retry succeeding faster.
  2. A measured rate drop with a pre-stated sample size, taken the same way before and after, using
    sliced queries.
  3. Time to first reply on a cold open improves, shown as a distribution rather than a median.

What is actually wrong

Measured over 72 hours on POST .../messages/stream:

200  n=138  p50=3000  p90=6000  max=10000 ms
503  n=61   p50=6000  p90=9000  max=13000 ms   (30.0 %)

The failures cost more than the successes. 87 % recover on the client's automatic retry about six
seconds later, so the user experiences roughly ten seconds to first token. 343 seconds wasted across
61 attempts.

The 503 is not a mystery, and the earlier framing of it as one was wrong. It is
SharedRuntimeCacheWarmingError, translated at shared-runtime-conversation.ts:1762-1777 to
conversation_cache_warming, then rewritten to shared_runtime_cache_warming by
conversation-coordinator.ts:251-255. That renaming is why grepping for warming in the logs found
nothing and made the branch look unreachable. Do not treat the absence of that string as evidence.

The source already documents this phenomenon. prewarm-shared-agent.ts:1-25 states that a fresh
shared agent's first turn is served entirely from caches, fail-closed, and
prewarm-shared-agent.ts:81 logs "[shared-runtime prewarm] leg failed; first turn falls back to warming 503s". That line, and canonical-scoped-stream.ts:209, are the diagnostic strings worth
grepping.

The fix most likely lives in the prewarm/keepwarm rail, which was missing from the earlier
analysis.
packages/cloud/api/v1/cron/shared-agent-keepwarm/route.ts:48-51 skips exactly the
identities that produce these 503s. Start there.

The one comparison worth making is not /bridge against /stream as populations — those differ
in composition. Split /bridge by rpc.method and compare /bridge message.send against /stream.
Those two run the same gates. If that gap is still large, the difference is real and diagnostic; if
it collapses, the earlier "60× gap" was an artifact of mixing methods.

Cold open is slow for a related but distinct reason. Across 15 real sessions, cold open to first
reply is p50 18.8 s, max 406 s, with a median of five sequential requests before the first reply and
13 of 15 containing at least one 503. Hibernation is refuted — idle makes it faster:
GET /api/conversations after >60 min idle is p50 20 ms against 110 ms after <10 s idle. What was
described in the original QA pass as a ~34 s "wake from sleep" is a serial startup burst against a
cold isolate plus a retry cycle. Reducing that burst is client-side sequencing work as much as
server work, and it compounds with GOAL C.

Definition of done

The three proofs, with the sample size stated in advance and all measurements taken from sliced
queries. Proof 1 in particular: show the first turn succeeding, not the second one arriving sooner.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    needs-triageOpen issue requires area, type, priority, and acceptance-mode triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions