Repository navigation
fix(managed-agent): the cancellation takeover never finishes against a replacement Broker #13171
Description
Activity
Independent source verification against main (728c13de) — root cause located and confirmed end-to-end.
1. Passive recovery reads status without acquiring. recoverHostedRuntimeTurn (packages/cli/src/serve/hosted-runtime-recovery.ts:247-254) branches on passive and, for pending.length > 0, loops broker.status(item.executionCallId) directly. The active branch (:255-269) and the no-pending re-attach branch (:371-388) both call broker.acquire() first; the passive branch is the only one that touches the Broker without adopting the Runtime Session. acquiredRuntime (:245) therefore stays false on the passive path, so the returned HostedRecoveryTurn (:432-444) tells the caller no lease was taken — the cancel route is left to release a Session this Broker never owned.
2. The real Broker refuses status/release for a session it never adopted. RuntimeBrokerService.requireSession (packages/sdk-java/runtime-broker/src/main/java/com/alibaba/qwen/code/runtimebroker/RuntimeBrokerService.java:2415-2433) reads sessions.get(runtimeId) and returns notFound("runtime_session_not_found", "Runtime Session is not active in this Broker process") (:2422-2424) when the replacement process has no live context. getExecution (:464-492) reaches this via requireReadySession (:487) for any non-terminal execution — i.e. the in-flight row → 404 → load answers 409 hosted_turn_recovery_required. release (:1015-1033) hits the same null-context path through releasedSession (:1035-1090), which throws unavailable("runtime_reconciliation_required", "Runtime Session is not active in this Broker process") (:1086-1088) unless the binding is LOST with stopped writers and no active executions (:1062-1085) — the settled-but-still-active row falls through to that throw → 503 managed_runtime_cancel_failed. Exactly the two rows in the table above.
3. The unit suite pins the contract the real Broker does not offer. hosted-harness-session.test.ts:4625-4627 stubs acquire to resolve with no-op, :4632 stubs release to resolve, :4827-4836 stubs status/cancel to succeed without any acquire. The passive test reports a parked execution passively and cancels the turn (:4824) then asserts expect(acquireSpy).not.toHaveBeenCalled() (:4842) — i.e. the test requires the passive path to skip acquire — and expect(release).toHaveBeenCalled() (:4873) — i.e. the cancel route releases a lease the Broker never granted. The comment at :4871-4872 states this explicitly. So the suite encodes a Broker contract (status/release succeed unconditionally) that the real RuntimeBrokerService only honours for a session it has adopted, which is precisely what an owner change prevents.
Reachability note — the issue's reachability analysis holds: ManagedAgentStore.insertCancelCommand refuses Workspace Sessions (409 workspace_unavailable) today, so the path is only reachable today by applying the status='CANCELLING' transition directly, then starting owner B. It goes live when Workspace cancel is opened. The cancellation takeover is otherwise contract-complete with no E2E mode, as the #13083 design states.
Fix direction. The candidate (acquire first in the passive branch when Runtime items exist, release-on-failure mirroring the active branch's :358-369 catch) is the right call: it makes the passive path's Broker interactions match the adoption contract that requireReadySession and release actually enforce, and the test assertion flip (acquireSpy called, plus an acquire stub on the three passive tests) corrects the pinned-but-wrong contract. The alternative — Broker-side: answer non-terminal status from the durable record and accept release on a never-adopted session — is broader and overlaps #13040's Broker-side reconciliation loop; the recovery-side fix is the minimal change that closes this path and keeps the Broker's adoption invariant intact.
Evidence re-derived from main at 728c13de (equals #13083 squash on the affected files). Posting as independent confirmation to unblock the candidate patch's review.
What happened?
The cancellation half of the Hosted Turn takeover added in #13083 cannot finish once the Turn's owner has changed. That half is the coordinator's passive takeover load followed by
POST /session/:id/managed-runtime/cancel.The replacement owner's Broker answers execution status and release only for a Runtime Session it has adopted, and the passive load never adopts one. So the Turn stays
CANCELLINGand the coordinator retries forever.I measured this on the packaged stack:
dist/cli.jsHosted Harness;It reproduces at
7ae1fa05, the last head of #13083. That head is identical, on the PR's 36 files, to the squash728c13denow onmain. It also reproduces at the earlier headb4e9d71b. Details are in round 4 of the real-environment verification.load: the Harness reads the execution status, the Broker answers404 runtime_session_not_found, and the load answers409 hosted_turn_recovery_requiredCANCELLINGload200; then the cancel route'sbroker.release()gets503 runtime_reconciliation_required("Runtime Session is not active in this Broker process"), and the route answers503 managed_runtime_cancel_failedCANCELLINGA lost cancel reply makes no difference, because the cancel route is never reached (in-flight) or never succeeds (settled).
Why:
recoverHostedRuntimeTurnwithpassive: truecallsbroker.status()without abroker.acquire()(packages/cli/src/serve/hosted-runtime-recovery.ts, passive branch).RuntimeBrokerService.getExecutiongoes throughrequireReadySession, which needs the Runtime Session to be active in this Broker process.releaseof a session the Broker never adopted is refused withruntime_reconciliation_required. This is the same refusal that was behind F2 in round 1 of feat(managed-agent): Hosted Turn takeover and G1 failover E2E #13083.statusandreleaseto succeed without an acquire.reports a parked execution passively and cancels the turn(hosted-harness-session.test.ts) even assertsexpect(acquireSpy).not.toHaveBeenCalled(). So the tests encode a Broker contract that the real Broker does not offer after an owner change.Reachability. Today the public API cannot put a Workspace Turn into
CANCELLING:ManagedAgentStore.insertCancelCommandrefuses Workspace Sessions with409 workspace_unavailable. The #13083 design also says the cancellation takeover "is implemented for contract completeness but has no E2E mode". The verification reached the path by applying the store's own cancel transition (status = 'CANCELLING') after owner A died, then starting owner B. The path becomes live as soon as Workspace cancel is opened.What did you expect to happen?
The coordinator's cancellation takeover should settle the Turn as cancelled:
Client information
Client Information
This is server-side, not a client bug:
qwen serve --profile hosted-harnesstogether withmanaged-agent-server. Verified on Linux (Ubuntu 24.04 container, arm64, JDK 21, Node 24.18, MySQL 8.0.46) at7ae1fa05, which equalsmainat728c13deon the affected files.Anything else we need to know?
Candidate fix:
candidate-r4.patch, about 30 lines inhosted-runtime-recovery.ts.acquireSpyassertion above and adds an acquire stub to the three passive recovery tests.On the same rig, with the candidate:
turn.cancelled, in 1.6 to 4.0 s. The cases are in-flight and settled, each with and without a lost cancel reply.mainat728c13de.The candidate relies on the Broker being able to adopt the dead owner's binding. That works with the durable local-process provisioner, which the takeover modes already require. With the default provisioner the acquire itself cannot adopt; that Broker-side loop is #13040.
The alternative is a Broker-side change: answer status from the durable record, and accept the release of a session the Broker has not adopted.
Evidence: the rig, raw results and Harness logs (Broker 404/503 lines) are in
pr13083/r4/. The scenarios areresults/s2/{h7,h8,h8c}-*db-cancel*, and the driver isharness/rig/s2-takeover.tswith faultsdb-cancelanddb-cancel-drop-reply.Related: #13083, #13040, #12380.
中文说明
发生了什么
#13083 加入的 Hosted 轮次接管里,cancel 这一半在轮次的 owner 变更之后走不完。这一半是协调器的被动接管 load,再加上
POST /session/:id/managed-runtime/cancel。替代方 owner 的 Broker 只对它已接管的 Runtime Session 回答执行状态和 release,而被动 load 从不接管。所以轮次一直停在
CANCELLING,协调器无限重试。在打包栈上实测:
dist/cli.jsHosted Harness;在 #13083 的最后一个 head
7ae1fa05上可以复现;该 head 在 PR 的 36 个文件上与main上的 squash 提交728c13de完全一致。在更早的 headb4e9d71b上也一样。细节见真实环境验证第四轮。load:Harness 读执行状态,Broker 返回404 runtime_session_not_found,load 返回409 hosted_turn_recovery_requiredCANCELLINGload返回 200;接着 cancel 路由的broker.release()得到503 runtime_reconciliation_required("Runtime Session is not active in this Broker process"),路由返回503 managed_runtime_cancel_failedCANCELLINGcancel 应答丢失与否没有区别:in-flight 时根本到不了 cancel 路由,已结算时 cancel 永远不成功。
原因:
recoverHostedRuntimeTurn在passive: true时直接调broker.status(),没有先broker.acquire()(packages/cli/src/serve/hosted-runtime-recovery.ts的 passive 分支)。RuntimeBrokerService.getExecution会经过requireReadySession,它要求该 Runtime Session 在本 Broker 进程中处于活动状态。release会被runtime_reconciliation_required拒绝。这就是 feat(managed-agent): Hosted Turn takeover and G1 failover E2E #13083 第一轮 F2 背后的同一个拒绝。status和release桩成不需要 acquire 也能成功。hosted-harness-session.test.ts里的reports a parked execution passively and cancels the turn甚至断言了expect(acquireSpy).not.toHaveBeenCalled()。也就是说,测试编码了一个真实 Broker 在 owner 变更后并不提供的契约。可达性。 目前公开 API 没法让 Workspace 轮次进入
CANCELLING:ManagedAgentStore.insertCancelCommand对 Workspace 会话返回409 workspace_unavailable。#13083 的设计文档也写明 cancel 接管「为契约完整而实现,没有 E2E 模式」。验证时是在 owner A 死后,直接执行 store 本来会做的取消转移(status = 'CANCELLING'),再启动 owner B 来触达的。一旦开放 Workspace 取消,这条路径就会真实出现。期望的行为
协调器的 cancel 接管应把轮次结算为已取消:
候选修复
candidate-r4.patch,hosted-runtime-recovery.ts约 30 行:acquireSpy断言反过来,并给三个被动恢复测试补了 acquire 桩。同一套装置上,加上候选之后:
turn.cancelled结束,耗时 1.6 到 4.0 秒。四个场景是:in-flight 和已结算,各自分有无 cancel 应答丢失。728c13de的main。候选修复依赖 Broker 能接管死掉的 owner 的绑定。用 durable 本地进程 provisioner 时可以做到,而接管模式本来就要求这个 provisioner。用默认 provisioner 时 acquire 本身就接管不了;那个 Broker 侧的循环见 #13040。
另一种做法是改 Broker 侧:未终结执行的状态从持久记录回答,并接受对未接管会话的 release。
证据: 装置、原始结果和 Harness 日志(含 Broker 404/503 行)在
pr13083/r4/。场景结果是results/s2/{h7,h8,h8c}-*db-cancel*,驱动脚本是harness/rig/s2-takeover.ts,故障名为db-cancel和db-cancel-drop-reply。关联:#13083、#13040、#12380。