Skip to content

Certify Shared-to-Dedicated continuity, ownership, and behavior parity #29921

Description

@standujar

Parent launch epic: #25020
Messaging analysis parent: #25076

Goal

The same external provider identity and conversation route a linked user to Personal Shared by default and to their active Dedicated runtime when explicitly selected, without bot replacement, conversation loss, silent activation, or tenant leakage.

In scope

Routing decisions; quote/consent; active, unavailable, starting, sleeping, revoked, and recovery states; DM and owner-bound group continuity; retry/fallback truth; provision/health/chat/sleep/wake/restore/chat; same conversation and provider identity.

Non-goals

Provider account setup; direct/group connector semantics; broad provisioning redesign; pricing redesign; production promotion.

Dependencies and safety gate

Front-door, #25025, Shared 1:1, and group Goals; #21594, #25044, #22934, #20615, #28500, #29622, #29653, #27096, and #25076. A composed live transition is forbidden until the explicit safety GO required by #25076 is recorded.

Acceptance criteria

  • One renewable Dedicated staging fixture completes provision → health → chat → sleep → wake/restore → chat.
  • Applicable provider conversations preserve bot identity, channel binding, conversation, and owner scope across routing changes.
  • No implicit lifecycle request occurs before explicit quote/consent.
  • Unavailable/starting/failure/fallback states are truthful and actionable.
  • Staging, tenant, private-memory, and production isolation pass.

Evidence

Provider receipt; Cloudflare router events; gateway logs; attested DB binding/lifecycle state; control-plane jobs; sandbox source SHA, heartbeat, runtime, sleep/wake, restore, and cleanup logs.

Execution contract

This is a bounded execution epic under #25020. Work validation-first: inspect the live staging product with Computer Use on desktop and mobile where applicable, then inspect the exact current code path before drawing conclusions. Use ui-ux-pro-max for customer-facing UX. Pin evidence to the served source/deployment SHA and correlate only applicable layers: browser/network, Cloudflare, gateway/Railway, PostgreSQL, control plane, sandbox/runtime, and provider receipts. Source, mocks, health endpoints, or a gateway 2xx do not prove live acceptance.

Keep secrets, OTPs, callback tokens, cookies, private fixture names, message content, and user identifiers out of GitHub and transcripts. Do not merge, deploy, promote, mutate production, or absorb another Goal's scope. Obtain exact just-in-time confirmation before any OTP/provider message, account/passkey/signature creation, OAuth approval, bot installation, group/webhook/provider mutation, identity reset/link/unlink/logout/revoke, mutating workflow dispatch, unattested DB query, or sandbox lifecycle action.

The executing agent must claim only this epic or one bounded child issue. A reproduced defect may produce one focused owning issue/PR; do not churn on adjacent code.

Required skills and tools

  • Load and follow browser:control-in-app-browser, computer-use:computer-use, and elizaos-ops; use ui-ux-pro-max for visible quote, consent, progress, and failure states.
  • Use cloudflare and configured Cloudflare logs/MCP; load wrangler before Wrangler. Use use-railway for gateways/services.
  • Use the elizaOS ops SSH/control-plane procedures for names-only environment checks and authorized in-place DB/control-plane/sandbox evidence. Do not use generic shell inspection that can expose credentials.
  • Load workers-best-practices before changing Cloudflare Worker code and find-docs/ctx7 for current infrastructure/provider contracts.
  • Use gh api, rg, Git, and repository-pinned validation commands for coordination and code.
  • Run review-security for routing, lifecycle, tenant isolation, DB, or sandbox changes and review-bugbot for every candidate PR.

Sprint 3 connector and routing qualification slice

Nubs is additive sprint integration lead; retain Stan's ownership and the existing execution/fixture gates. Nominal 4 person-days (3–5 plus concrete remediation), High complexity. This allocation covers the integrated maintained-ingress proof using #29919 and #29920, not those entire epics or #22934's broad program counted again.

Explicit product direction: once the user's selected Dedicated is active and entitled, it owns that user's turns across every maintained gateway; do not also execute Shared. Two users with Dedicated agents on the same gateway identity must reach their distinct account-bound Dedicated targets. Test one Shared plus one Dedicated, two Dedicated users, group owner/guest, same account on multiple linked connectors, stale cache, duplicate/reordered ingress, target change mid-turn, sleeping/starting/failed instance, disconnect/reconnect and late egress. Require one route and one attributable reply.

Every maintained connector must prove central Cloud sign-in → correct account link → return to original provider conversation → repeat reply, including hosted Blooio/iMessage. Native macOS Messages transport is a separate target and must not be mislabeled Cloud sign-in. Use actual registered support for Discord, Telegram, hosted iMessage/Blooio, WhatsApp, X and applicable SMS/voice/native transports; unsupported paths stay explicitly classified. Free DB/API capabilities, media allowance outcomes and upsell/sign-in continuations must work on these ingresses, not just web. Do not send private account-recovery state into an owner/guest group by default; use the authorized private continuation.

Integrate #25146's UPDATED isolated Shared fallback: no Dedicated memory becomes accessible to Shared during bad standing. Recovery resumes Dedicated only after authoritative entitlement/health and scoped Shared-journal reconciliation. Do not reuse #25146's superseded import-on-fallback requirement.

Evidence: real per-provider inbound/outbound/readback receipts, route account/runtime identity and transition generation, canonical DB binding/readback and live model trajectories. HTTP 200, a source test or UI screenshot alone is insufficient. Environment/account/fixture blockers remain named with owning issues. No live messages, link changes or provisioning were executed in this planning pass.

Project 19 Sprint 3 (October 3–16, 2026): scope candidate, not a capacity-committed completion date. Planning estimate: 4 person-days for this slice. Preserve existing ownership; @NubsCarson is the sprint integration lead. Full scope exceeds two-person sprint capacity; sequencing and remaining policy decisions are in the project README.

Critical qualification plan — 2026-09-06

Source baseline: 8fea4f94c29c7712ce3caf303e3350150622d048. This is source/issue research, not validation-first live execution. No provider message, link/account operation, sandbox lifecycle, deployment or DB query was performed.

Current implementation and evidence limits

Bounded execution and remediation sequence

  1. Readiness and owner reconciliation. Preserve Stan's ownership and Nubs's additive integration lead. Rebaseline Provision and certify staging account and managed-provider identities #25025 identities/fixtures, Isolate Cloud database authority and prove environment and tenant recovery #21594 attested DB authority, [P0][staging] Restore a renewable Dedicated fleet: zero usable staging fixture #25044 renewable Dedicated lifecycle, Certify reliable 1:1 messaging through maintained provider identities #29919 ingress, Certify owner-bound group messaging and complete delivery receipts #29920 groups and [P0][staging][messaging] Certify shared-bot onboarding, group ownership, and Shared/Dedicated routing #25076 safety status. Record served source/deployment and exact source of gateways/sandbox separately. Before live composed transition, obtain the existing explicit safety GO and action-specific just-in-time confirmations required by this issue; planning adds no authorization. Do safe source/harness work while fixtures are blocked.
  2. Create a per-provider qualification matrix. Enumerate registered maintained Discord, Telegram, hosted Blooio/iMessage, WhatsApp, X and applicable SMS/voice/native routes; mark supported, unsupported, missing fixture, failed or verified with evidence. Keep native macOS Messages distinct from hosted Blooio Cloud sign-in. For each supported lane identify central login continuation, canonical account binding, ingress router, target authority, delivery dedupe and egress acknowledgement. Reuse existing owner artifacts instead of rebuilding provider setup.
  3. Prove the same-provider base journey. Using an authorized controlled fixture, observe unknown sender → central Cloud sign-in → account link → return to original provider conversation → first and repeat Shared reply. Record provider identity/channel/conversation binding throughout. Then explicit quote/consent selects the entitled Dedicated and one existing lifecycle command provisions/health-checks it. No implicit provisioning or new provider bot is allowed. Validate actual provider outbound/readback, not 202/HTTP 200.
  4. Prove routing ownership under concurrency. Test one Shared plus one Dedicated user and two distinct Dedicated users on the same gateway identity; same account across linked connectors; duplicate/reordered ingress; target changes mid-turn; stale caches; disconnect/reconnect and late replies. Require one authoritative route generation and one attributable reply, with no simultaneous Shared execution for an active entitled selected Dedicated. Test group owner/guest boundaries only on the approved group lanes and route sensitive recovery actions through authorized private continuation.
  5. Exercise lifecycle and fallback. Complete provision → health → chat → sleep → wake/restore → chat with the same provider identity/conversation and current Dedicated data. Check starting/unavailable/blocked states against actual worker/job/runtime state. Consume [cloud][billing][shared] Add reversible entitlement-driven Dedicated-to-Shared recovery with pay action #25146 for confirmed entitlement withdrawal: isolated Shared fallback, zero protected Dedicated retrieval, eligible free operations and truthful account recovery. Resume Dedicated only after entitlement/health plus complete scoped Shared-interval reconciliation. Payment redirects or incidental cash balance are not routing authority. Include deletion and canceled recovery.
  6. Remediate only demonstrated owning failures. When live evidence shows mismatch, pin first failing account/route/lifecycle/delivery transition and exact SHA. Reproduce with the narrow real-runtime/DB regression, then change the canonical owning seam; do not add provider-local routing fallbacks. Route identity/ingress/group/provisioning defects to their existing owners. Run independent routing/security review and applicable package/root checks; requalify affected providers at the final served source.
  7. Publish complete receipts and disposition. Join provider receipts, Worker route events, gateway logs, attested DB binding/transition generation, lifecycle jobs and sandbox source/heartbeat/restore logs. Inspect desktop/mobile quote/consent/progress/failure UI and complete live-model trajectories, including free API/media allowance and sign-in/upsell continuation. Use the canonical verified evidence bundle with private fixture data removed from public projections. Keep unavailable source, fixture, account or policy evidence explicitly blocked.

Acceptance/estimate boundary: four nominal days (3–5 plus remediation) cover integrated maintained-ingress qualification after identity, group, billing/fallback and renewable fixture prerequisites. It does not include those epics' missing implementation or provider setup. A verified provider lane cannot certify an untested lane; a successful Shared turn cannot certify the composed Dedicated transition. The issue remains open until the required applicable matrix and isolation/lifecycle receipts pass, with unsupported paths explicitly justified.

Independent deep review: #29921 — same-provider cutover/recovery certification

Source b676280384b15df54204276e34bb8740db26936f; live issue/comments and current routing, bridge, import and identity consumers reviewed. This adds a concrete recovery failure to the qualification matrix. No connector message, account linking, billing, provisioning or live transition was performed.

P1 — Failed nonempty history restoration is replaced by an empty import and returned as success (reproduced)

On Bridge returned HTTP 404, deliverPersonalTextMessage fetches Shared history and tries to import it. If a nonempty full import returns no receipt—or some historical messages lack source IDs—it calls importCanonicalConversation again with an empty array. A successful empty receipt permits retrying the message; an ordinary reply then returns success:true. The personal-shared connector route duplicates this behavior for its maintained ingress paths and group-specific room.

Executed bun /tmp/guardian-project19-deep-review/probe-29921.ts against the source-extracted unchanged complete delivery helper, with controlled bridge/history/import collaborators:

Input state Observed operations Final result
One historical message; full import rejects bridge404 → import(1) fails → import(0) succeeds → bridge retry success:true
Two historical messages, one lacks ID bridge404 → import(0) succeeds → bridge retry success:true
Genuinely empty source history bridge404 → import(0) → bridge retry success:true — valid control
Complete nonempty history bridge404 → import(1) succeeds → bridge retry success:true — valid control

The probe proves success can bypass the caller's original nonempty restoration requirement. It does not prove data was deleted, a real remote import failed, or a provider received a reply.

Crucial counterevidence: the importer correctly validates inserted+skipped and modern/legacy receipts against the supplied message count. Its empty receipt is truthful for zero submitted messages. The caller silently changed the expected source set; adding another complete:true check there would not fix this defect. The agent import endpoint supports canonical import, so a legitimate empty new session must remain possible.

Required plan amendment

  1. Resolve a typed repair decision distinguishing a genuinely new/empty conversation from required-history restoration. Capture immutable expected source IDs/count and provenance before import. Preserve the entire selected history; missing IDs require an explicit migration/reconciliation outcome, not omission or fabricated IDs without a migration contract.
  2. Remove empty fallback when a nonempty source requires restoration. Return a typed recoverable unavailable/incomplete state and keep the same input message identity pending. A failure may have partially imported data, so retry the same canonical import idempotently and reconcile actual receipt/state instead of creating another conversation.
  3. Consolidate the duplicated repair policy into the canonical delivery/import seam while retaining provider/group identity and the existing private-vs-group room separation. The generic helper is currently called by the X personal-DM route; it is not evidence that all providers already share this entire implementation. The other connector route needs the same repair.
  4. Bind completion evidence to the original required source set and target conversation, not only the latest import call's count. For [cloud][billing][shared] Add reversible entitlement-driven Dedicated-to-Shared recovery with pay action #25146 recovery, substitute the explicitly authorized fallback interval and generation; do not select the old pre-upgrade archive or permit Dedicated-to-Shared import.
  5. Add actual DB+agent import tests for full import rejection, partial import then retry, mixed/missing source IDs, process restart after import-before-message retry, and a genuinely empty new conversation. Preserve same input message ID on retry and prove the resulting model request includes every required historical message exactly once before substantive generation. Then run the authorized same-provider journey with a distinctive prior fact; receipt/readback and the complete model context must agree.

Examined protections and unresolved boundaries

Dedicated target selection joins canonical account-owned authority, validates the cutover schema and quarantines unverified markers. Sleeping/error targets remain Dedicated rather than splitting traffic to Shared; preparation returns explicit blocked/starting/unavailable. X prefixes its incoming ID with x-dm: before the helper, so this review does not allege an unscoped X message-ID collision. The bridge forwards clientMessageId and the agent conversation path has durable admission/replay machinery; don't add blind retries as if no idempotency foundation existed. Confirm provider egress outcomes separately: helper reply success is runtime generation, not recipient delivery.

The broader matrix still needs actual two-user routing, same-provider group/DM isolation, unlink/relink, stale-generation races, sleep/wake and entitlement fallback receipts under #25076's established live gates. This review neither grants that live GO nor certifies the untouched providers. The source-extracted probe passed; full runtime/DB/evidence fixtures were not executed.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions