Skip to content

[cloud][billing][shared] Add reversible entitlement-driven Dedicated-to-Shared recovery with pay action #25146

Description

@lalalune

Product contract — revised for Sprint 3

When confirmed server-side billing/entitlement policy withdraws Dedicated access, the user may continue eligible free operations on Personal Shared. Shared must have no access to Dedicated memory during fallback. This requirement supersedes the former import-before-fallback design. Preserve Dedicated data for entitled recovery; do not copy, retrieve, summarize, embed, or include it in Shared prompts or fallback context. A shared schema or database does not grant cross-tier data access.

Shared receives a typed, minimal account provider: access state, confirmed reason category, safe deadline if applicable, available free capabilities, and an authorized sign-in/billing recovery action. Say plainly that Dedicated memory is unavailable until billing is restored. Do not expose card details, raw Stripe objects, secrets, or unrelated account data. Keep normal free Calendar/Notes and other eligible capability state separate from protected Dedicated memory; restoration of a schema pointer must not bypass that boundary.

Existing architecture and remaining gap

The current personal target resolver deliberately retains Dedicated authority through sleep/error/restart and completed cutover markers, without a billing-loss state machine. Reuse the canonical identity, route, entitlement and durable command authorities; do not add another gateway router or billing scheduler. Dependencies: #30638, #23091, #23093, #23094, #23095, #23096, #23098; same-bot certification #29921 and ingress/group contracts #29919/#29920.

Required behavior

  • One typed durable transition owner with revision/generation fencing: dedicated_active → fallback_pending → shared_active → recovery_pending → dedicated_active, plus explicit waiting/failed/unavailable outcomes. Confirm current account/tenant/owner entitlement at every effect and final route commit.
  • Distinguish confirmed nonpayment/expired access from low credit, trial expiration, grace/dunning, canceled-at-period-end, provider outage, webhook lag, sleeping/starting Dedicated and unknown billing state. Product policy must define which condition withdraws entitlement; never equate every transient error or zero incidental cash balance to subscription loss.
  • Finish, cancel or honestly fail in-flight work under one request owner. Route new requests to one active destination; retries and stale workers cannot generate duplicate replies, writes, charges or instance starts.
  • Before Shared activation, establish a separate scoped fallback conversation/journal and revoke Shared access to Dedicated context. There is no Dedicated-to-Shared memory import step. Pre-upgrade archives or cached provider context cannot be used to reconstruct protected Dedicated history.
  • Persist the Shared interval completely and with provenance. After confirmed payment/entitlement recovery, Dedicated may idempotently reconcile that explicitly scoped interval, verify health/current authority and then own routing again. Maintain data lineage and user-visible continuity without exposing Dedicated data while access is suspended.
  • Recovery links are signed/expiring and require appropriate interactive identity and role at use. They return the user to the original authorized conversation after successful billing reconciliation; a checkout redirect alone cannot restore access.
  • Dedupe transition, warning/deadline, payment-action and recovery notifications by durable generation. Billing/system notices remain deliverable without a paid workflow subscription.

Acceptance and edge cases

  • A confirmed expired/unpaid account falls to Shared, receives conversational account-aware recovery help, and can perform eligible free operations. Asking Shared about Dedicated-only prior work yields accurate unavailable state and zero Dedicated retrieval.
  • Subscription active but usage/media allowance exhausted is not mislabeled unpaid. Canceled-at-period-end retains entitled access until the approved boundary. Trial, grace, late invoice, revoked account, refund/dispute and provider-unknown policies are explicit.
  • Two paid users on one gateway each route to their own Dedicated, never Shared or the other user's instance. Mixed Shared/Dedicated users, linked connectors, owner-bound groups, unlink/relink, owner changes and stale routing caches preserve authority and reply destination.
  • Payment success/replay/reordered webhook, failure during provisioning, crash before/after route commit, suspended/failed instance, repeated recovery, cancellation during recovery and deletion all converge without duplicate effects or premature memory access.
  • Memory/state references retain complete values; no lossy compaction or implicit recent window. A failed reconciliation is an explicit recoverable state, not fabricated success.

Verification

Real PostgreSQL/PGlite transition, restart and concurrency contracts; real signed Stripe sandbox lifecycle; dedicated instance health/route readback; actual ingress/egress receipts; live-model fallback and recovered trajectories; controlled account fixtures proving memory remains unavailable throughout fallback. Include desktop/mobile recordings and inspected canonical evidence. Fixtures/mocks alone do not prove live connector routing. Keep private fixture data out of public evidence. No production payment/resource change or live connector message was performed in planning.

Sprint 3 scope allocation

Project: https://github.com/orgs/elizaOS/projects/19 . Shaw lead; high complexity; 6 person-days nominal, 5–8 estimated. Candidate in October 3–16 scope, not a deadline commitment; full Sprint 3 is over capacity. Product decisions on grace, retention, resource funding and recovery are prerequisites, not values inferred here. This replaces the incompatible old fallback-memory acceptance; historical issue revisions remain available.

Project 19 Sprint 3 (October 3–16, 2026): scope candidate, not a capacity-committed completion date. Planning estimate: 6 person-days for this slice. Preserve existing ownership; @lalalune is the sprint integration lead. Full scope exceeds two-person sprint capacity; sequencing and remaining policy decisions are in the project README.

Critical implementation audit — 2026-09-06

Source: 8fea4f94c29c7712ce3caf303e3350150622d048. Source/live issue research only; no account, message, billing, infrastructure or data migration operation occurred.

Existing routing and incompatible historical assumptions

Personal target resolution retains a completed server-owned Dedicated cutover through sleep/error/restart. That prevents split histories; it is not an entitlement-driven fallback state machine. Preparation returns ready for running, checks the legacy credit gate before idempotent sleep/stop recovery, and otherwise returns explicit blocked/starting/unavailable. The connector route returns 402/503 for those outcomes rather than activating Shared. Personal text delivery is another consumer of the same target/preparation seam. Both must adopt one authority decision.

Existing cutover coordination can seal Shared admission, wait for an in-flight turn, snapshot and commit/release an exact lease. This is useful concurrency infrastructure, but simply reopening the original Shared identity/archive would violate the revised memory boundary. The historical comment claiming “Shared history reconciliation” is not acceptance of import-on-fallback: Dedicated → Shared transfer is forbidden during fallback, including pre-upgrade history that belongs to the protected Dedicated context. Only the newly scoped Shared interval can move in the recovery direction.

Detailed implementation plan

  1. Define provider-neutral eligibility from feat(billing): generic app subscriptions with independent seven-day trials #30638/[cloud][billing] Implement Stripe subscription checkout, Customer Portal, changes, and webhook reconciliation #23093/[cloud][billing] Derive every enforced rate and resource quota from the subscription entitlement projection #23094 with account/owner, current entitlement revision, decision reason and availability. Product owners define trial/grace/dunning/access/retention boundaries; Define and implement payment-reversal, entitlement, and creator-payout policy #22930 owns reversal/refund/hold restoration. Low cash, active subscription allowance exhaustion, sleeping compute, provider outage and unknown webhook state are not confirmed entitlement withdrawal. Preserve the existing Dedicated route on unknown/transient states with honest unavailability.
  2. Extend the canonical durable routing/cutover authority with a revision-fenced transition generation and explicit pending/active/recovering/failed outcomes. Use the existing command/lifecycle/queue owners; no new gateway router, billing poller or competing scheduler. Serialize requests against one active destination. Recheck account ownership, deletion state and entitlement at every external effect and final route CAS.
  3. On confirmed withdrawal, drain/cancel/honestly fail the current request under one delivery owner and establish a new scoped fallback journal/context identity before admitting Shared. Deny Shared reads from Dedicated stores, pre-upgrade archives, restored pointers and cached providers. Keep eligible free calendar/notes records accessible through their own authorized domain contract, not by importing Dedicated memory. This is an enforced repository/provider access boundary, not just a prompt instruction.
  4. Add the minimal account-state provider with typed access/reason/deadline/free-capabilities and authorized recovery action; omit private payment/provider details. Produce a plain truthful memory-unavailable reply when asked about protected Dedicated history. Bind recovery continuation to interactive account authorization and original conversation; use private continuation for group account recovery. Dedupe warning/deadline/action notices by durable transition generation; notices must remain deliverable without a paid workflow entitlement.
  5. Recover only after authoritative paid/trial entitlement restoration and Dedicated health. Seal the fallback interval only, import/reconcile complete interval data with provenance and idempotent receipts, then commit routing back to the same authorized Dedicated. A successful checkout redirect is insufficient. Crash/retry at each stage must resume the same command/instance; failed reconciliation remains explicit and must not expose Dedicated memory through Shared. Never summarize, compact or silently omit history.
  6. Extend target/delivery/coordinator tests and real DB concurrent transition/restart coverage: repeated withdrawal/recovery, entitlement changes during import, stale workers/cache, webhook reorder, in-flight late replies, account deletion, owner/link changes and two users sharing one gateway identity. Use distinct marker data to prove zero Dedicated retrieval while fallback is active and complete eligible Shared interval recovery afterward.
  7. Under Certify Shared-to-Dedicated continuity, ownership, and behavior parity #29921/[P0][staging][messaging] Certify shared-bot onboarding, group ownership, and Shared/Dedicated routing #25076's existing live gates, run signed sandbox billing transitions with actual Dedicated health/route readback and provider ingress/egress receipts. Inspect Shared and recovered live-model requests plus authorized DB artifacts, desktop/mobile recordings and the verified evidence bundle. Run affected shared/API/core tests/typecheck/lint/root verify, app audit for UI changes, and independent security review of the memory/tenant boundary. Public evidence excludes private fixture content.

Scope/capacity: six nominal days (5–8) cover this new reversible authority and isolated-memory workflow only once generic billing, policy and renewable fixture dependencies are ready. Existing upgrade helpers do not eliminate the need for new transition and access proofs. #29919/#29920 own ingress/group semantics; #29921 certifies integrated routing; do not absorb their epics. A search for PRs mentioning #25146 found none, which is not proof no active implementation owner exists—coordinate with the existing claimant before writing.

Independent deep review: #25146

Revision b676280384b15df54204276e34bb8740db26936f. Reviewed live revised contract, target authority, delivery preparation and both consumers, Personal Shared identity, runtime execution, REST history and cutover coordination. The reversible entitlement workflow remains missing; the deeper findings below identify specific integration constraints rather than claiming an existing fallback was exercised.

P1 implementation constraint — Isolate the fallback journal without inventing a new account authority

The proposed “new scoped fallback journal/context identity” needs two distinct identities. Personal Shared authority derives the exact agent ID from organization+user and verifies equality. Runtime execution grants the authenticated personal-user flag and personal media/reminder ports only for canonical identity plus platform funding. Appending a fallback generation to agent.id to obtain a clean history would fail that authority check and also change todo/reminder scopes. Conversely, retaining both old agent and old room reopens the old journal the revised product contract prohibits.

Keep canonical account authority stable and introduce a separately authorized fallback conversation/journal generation, resolved server-side from the entitlement transition. Coordinator addressing already accepts independent agent/room keys; use this seam deliberately. Current personal text delivery and its Shared branch pass the canonical agent ID as room. Update all ingress, REST history, model/cache, reminder provenance and recovery consumers through one typed active-journal mapping. The current REST conversation list/create/update still derive a single conversation from agent identity; changing only the connector writer leaves clients reading a different history.

Acceptance: same authenticated account/free capabilities across fallback, a distinct empty authorized fallback journal at activation, and no old-journal data in the actual model/provider request. Repeated fallback generations cannot reuse an earlier suspended interval inadvertently. Separately prove permitted free domain records survive; isolating conversation history must not silently erase todos/calendar capability state.

P1 implementation constraint — Readiness is not current entitlement or an effect authorization fence

preparePersonalDedicatedDelivery returns ready immediately for running. Sleeping/stopped checks legacy credit, awaits worker health, and queues a wake/resume with account IDs but no entitlement revision. Executed the unchanged source-extracted helper in /tmp/guardian-project19-deep-review/probe-25146.ts: running produced ready with zero gate calls; a controlled entitlement-state change during the health await still reached enqueue. This proves helper control flow, not actual unauthorized compute: the controlled entitlement variable is an external scenario marker, and downstream workers may impose additional policy.

The new transition cannot be installed only inside the sleeping branch or only in billing webhooks. Produce an authoritative typed entitlement/route decision before both running and sleeping delivery; fence lifecycle jobs and final route commit with its revision, and revalidate after awaited work. Unknown billing/worker state must preserve explicit unavailability, not trigger fallback. Run actual worker/DB races with withdrawal during health check, provisioning, interval import and final commit. Assert stale jobs cannot reactivate routing or create additional instances after a newer generation wins.

P2 integration constraint — Repair imports need the active interval, not the historic canonical room

Both Dedicated delivery consumers contain a bridge-404 recovery import from coordinateSharedHistory(agent.id, originalConversationId). The deeper #29921 review identifies their empty-import fallback after a failed full import. Even after repairing that completeness defect, this old-room repair path cannot be reused for entitlement recovery without an authorized interval selection. It would select the pre-upgrade journal instead of the newly scoped fallback interval. Reconciliation must carry journal generation, source interval bounds/provenance, complete source count/digest and a receipt bound to the target and transition. Reuse actual import machinery, not its old identity assumptions; do not transfer any Dedicated history into Shared.

Existing protections to preserve and precise remaining proof

Target selection checks account/organization ownership, server-owned cutover authority/version and quarantines unverified markers. It deliberately retains Dedicated authority through sleep/error. The cutover coordinator serializes admission and seals an exact snapshot under a token. These are real protections; removing authority or releasing an old seal merely to route Shared is not an acceptable shortcut.

No observed production entitlement withdrawal, memory leak, billing provider run, DB transition or infrastructure mutation is claimed. The missing workflow's acceptance requires actual read-boundary isolation and revision-fenced effects, not a prompt that says memory is unavailable. Extend the first plan with the identity/journal split, all-consumer mapping and current-entitlement fences above before estimating implementation closure. Bun helper probe passed; full DB/runtime tests require the established authorized fixtures and were not run.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

needs-triageOpen issue requires area, type, priority, and acceptance-mode triage

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions