Skip to content

fix(managed-agent): the cancellation takeover never finishes against a replacement Broker #13171

Description

@wenshao

What happened?

The cancellation half of the Hosted Turn takeover added in #13083 cannot finish once the Turn's owner has changed. That half is the coordinator's passive takeover load followed by POST /session/:id/managed-runtime/cancel.

The replacement owner's Broker answers execution status and release only for a Runtime Session it has adopted, and the passive load never adopts one. So the Turn stays CANCELLING and the coordinator retries forever.

I measured this on the packaged stack:

  • Spring fat jar, dist/cli.js Hosted Harness;
  • Runtime Broker with durable local workers;
  • MySQL, on Linux.

It reproduces at 7ae1fa05, the last head of #13083. That head is identical, on the PR's 36 files, to the squash 728c13de now on main. It also reproduces at the earlier head b4e9d71b. Details are in round 4 of the real-environment verification.

When owner A died, the parked execution was What owner B's passive takeover gets Result
in flight (not settled) load: the Harness reads the execution status, the Broker answers 404 runtime_session_not_found, and the load answers 409 hosted_turn_recovery_required 6 retries in 90 s; the Turn stays CANCELLING
already settled load 200; then the cancel route's broker.release() gets 503 runtime_reconciliation_required ("Runtime Session is not active in this Broker process"), and the route answers 503 managed_runtime_cancel_failed 6 retries in 90 s; the Turn stays CANCELLING

A lost cancel reply makes no difference, because the cancel route is never reached (in-flight) or never succeeds (settled).

Why:

  • recoverHostedRuntimeTurn with passive: true calls broker.status() without a broker.acquire() (packages/cli/src/serve/hosted-runtime-recovery.ts, passive branch).
  • For a non-terminal execution, RuntimeBrokerService.getExecution goes through requireReadySession, which needs the Runtime Session to be active in this Broker process.
  • A release of a session the Broker never adopted is refused with runtime_reconciliation_required. This is the same refusal that was behind F2 in round 1 of feat(managed-agent): Hosted Turn takeover and G1 failover E2E #13083.
  • The unit suites stub status and release to succeed without an acquire. reports a parked execution passively and cancels the turn (hosted-harness-session.test.ts) even asserts expect(acquireSpy).not.toHaveBeenCalled(). So the tests encode a Broker contract that the real Broker does not offer after an owner change.

Reachability. Today the public API cannot put a Workspace Turn into CANCELLING: ManagedAgentStore.insertCancelCommand refuses Workspace Sessions with 409 workspace_unavailable. The #13083 design also says the cancellation takeover "is implemented for contract completeness but has no E2E mode". The verification reached the path by applying the store's own cancel transition (status = 'CANCELLING') after owner A died, then starting owner B. The path becomes live as soon as Workspace cancel is opened.

What did you expect to happen?

The coordinator's cancellation takeover should settle the Turn as cancelled:

  • parked executions are stopped, or confirmed settled, without any dispatch;
  • the Workspace lease is released, so another Session can use the Workspace;
  • a replayed cancel is answered at its admission watermark.

Client information

Client Information

This is server-side, not a client bug: qwen serve --profile hosted-harness together with managed-agent-server. Verified on Linux (Ubuntu 24.04 container, arm64, JDK 21, Node 24.18, MySQL 8.0.46) at 7ae1fa05, which equals main at 728c13de on the affected files.

Anything else we need to know?

Candidate fix: candidate-r4.patch, about 30 lines in hosted-runtime-recovery.ts.

  • When a passive load has Runtime items, it acquires the Runtime Session first. Acquiring dispatches nothing.
  • If the acquire or a status read fails, it hands the lease back.
  • It flips the acquireSpy assertion above and adds an acquire stub to the three passive recovery tests.

On the same rig, with the candidate:

  • All four cases end turn.cancelled, in 1.6 to 4.0 s. The cases are in-flight and settled, each with and without a lost cancel reply.
  • In flight, the tool never runs, and undo has nothing to restore.
  • The lease is released, and a second Session on the same Workspace then runs.
  • A lost cancel reply is replayed at the same watermark (13 → 13, 18 → 18).
  • The drive path is unchanged: two-tool takeovers complete, and the feat(managed-agent): Hosted Turn takeover and G1 failover E2E #13083 runner passes 3/3 in both modes.
  • The changed CLI test files pass.
  • It applies cleanly to main at 728c13de.

The candidate relies on the Broker being able to adopt the dead owner's binding. That works with the durable local-process provisioner, which the takeover modes already require. With the default provisioner the acquire itself cannot adopt; that Broker-side loop is #13040.

The alternative is a Broker-side change: answer status from the durable record, and accept the release of a session the Broker has not adopted.

Evidence: the rig, raw results and Harness logs (Broker 404/503 lines) are in pr13083/r4/. The scenarios are results/s2/{h7,h8,h8c}-*db-cancel*, and the driver is harness/rig/s2-takeover.ts with faults db-cancel and db-cancel-drop-reply.

Related: #13083, #13040, #12380.

中文说明

发生了什么

#13083 加入的 Hosted 轮次接管里,cancel 这一半在轮次的 owner 变更之后走不完。这一半是协调器的被动接管 load,再加上 POST /session/:id/managed-runtime/cancel。

替代方 owner 的 Broker 只对它已接管的 Runtime Session 回答执行状态和 release,而被动 load 从不接管。所以轮次一直停在 CANCELLING,协调器无限重试。

在打包栈上实测:

  • Spring fat jar,dist/cli.js Hosted Harness;
  • 带 durable 本地 Worker 的 Runtime Broker;
  • MySQL,Linux。

在 #13083 的最后一个 head 7ae1fa05 上可以复现;该 head 在 PR 的 36 个文件上与 main 上的 squash 提交 728c13de 完全一致。在更早的 head b4e9d71b 上也一样。细节见真实环境验证第四轮。

owner A 死时挂起的执行 owner B 的被动接管得到什么 结果
正在执行(未结算) load:Harness 读执行状态,Broker 返回 404 runtime_session_not_found,load 返回 409 hosted_turn_recovery_required 90 秒内重试 6 次,轮次一直停在 CANCELLING
已结算 load 返回 200;接着 cancel 路由的 broker.release() 得到 503 runtime_reconciliation_required("Runtime Session is not active in this Broker process"),路由返回 503 managed_runtime_cancel_failed 90 秒内重试 6 次,轮次一直停在 CANCELLING

cancel 应答丢失与否没有区别:in-flight 时根本到不了 cancel 路由,已结算时 cancel 永远不成功。

原因:

  • recoverHostedRuntimeTurn 在 passive: true 时直接调 broker.status(),没有先 broker.acquire()(packages/cli/src/serve/hosted-runtime-recovery.ts 的 passive 分支)。
  • 对未终结的执行,RuntimeBrokerService.getExecution 会经过 requireReadySession,它要求该 Runtime Session 在本 Broker 进程中处于活动状态。
  • Broker 从未接管的会话,release 会被 runtime_reconciliation_required 拒绝。这就是 feat(managed-agent): Hosted Turn takeover and G1 failover E2E #13083 第一轮 F2 背后的同一个拒绝。
  • 单测把 status 和 release 桩成不需要 acquire 也能成功。hosted-harness-session.test.ts 里的 reports a parked execution passively and cancels the turn 甚至断言了 expect(acquireSpy).not.toHaveBeenCalled()。也就是说,测试编码了一个真实 Broker 在 owner 变更后并不提供的契约。

可达性。 目前公开 API 没法让 Workspace 轮次进入 CANCELLING:ManagedAgentStore.insertCancelCommand 对 Workspace 会话返回 409 workspace_unavailable。#13083 的设计文档也写明 cancel 接管「为契约完整而实现,没有 E2E 模式」。验证时是在 owner A 死后,直接执行 store 本来会做的取消转移(status = 'CANCELLING'),再启动 owner B 来触达的。一旦开放 Workspace 取消,这条路径就会真实出现。

期望的行为

协调器的 cancel 接管应把轮次结算为已取消:

  • 挂起的执行被停住(或确认已结算),全程不派发;
  • Workspace 租约被释放,其他会话可以使用该 Workspace;
  • 重放的 cancel 按其 admission 水位应答。

候选修复

candidate-r4.patch,hosted-runtime-recovery.ts 约 30 行:

  • 被动 load 有 Runtime 项时,先 acquire 这个 Runtime Session。acquire 不会派发任何东西。
  • acquire 或读状态失败时,把 lease 交还。
  • 测试方面,把上面那条 acquireSpy 断言反过来,并给三个被动恢复测试补了 acquire 桩。

同一套装置上,加上候选之后:

  • 四个场景都以 turn.cancelled 结束,耗时 1.6 到 4.0 秒。四个场景是:in-flight 和已结算,各自分有无 cancel 应答丢失。
  • in-flight 时工具从未执行,撤销也没有可恢复的内容。
  • lease 已释放,同一 Workspace 上的第二个会话随后能运行。
  • cancel 应答丢失后,重放拿到同一个水位(13 → 13、18 → 18)。
  • 驱动路径不变:两个工具的接管能完成,feat(managed-agent): Hosted Turn takeover and G1 failover E2E #13083 的 runner 两个模式都是 3/3。
  • 改动的 CLI 测试文件全部通过。
  • 可以直接应用到 728c13de 的 main。

候选修复依赖 Broker 能接管死掉的 owner 的绑定。用 durable 本地进程 provisioner 时可以做到,而接管模式本来就要求这个 provisioner。用默认 provisioner 时 acquire 本身就接管不了;那个 Broker 侧的循环见 #13040。

另一种做法是改 Broker 侧:未终结执行的状态从持久记录回答,并接受对未接管会话的 release。

证据: 装置、原始结果和 Harness 日志(含 Broker 404/503 行)在 pr13083/r4/。场景结果是 results/s2/{h7,h8,h8c}-*db-cancel*,驱动脚本是 harness/rig/s2-takeover.ts,故障名为 db-cancel 和 db-cancel-drop-reply。

关联:#13083、#13040、#12380。

Activity

doudouOUC commented on Oct 1, 2026

@doudouOUC
Collaborator

Independent source verification against main (728c13de) — root cause located and confirmed end-to-end.

1. Passive recovery reads status without acquiring. recoverHostedRuntimeTurn (packages/cli/src/serve/hosted-runtime-recovery.ts:247-254) branches on passive and, for pending.length > 0, loops broker.status(item.executionCallId) directly. The active branch (:255-269) and the no-pending re-attach branch (:371-388) both call broker.acquire() first; the passive branch is the only one that touches the Broker without adopting the Runtime Session. acquiredRuntime (:245) therefore stays false on the passive path, so the returned HostedRecoveryTurn (:432-444) tells the caller no lease was taken — the cancel route is left to release a Session this Broker never owned.

2. The real Broker refuses status/release for a session it never adopted. RuntimeBrokerService.requireSession (packages/sdk-java/runtime-broker/src/main/java/com/alibaba/qwen/code/runtimebroker/RuntimeBrokerService.java:2415-2433) reads sessions.get(runtimeId) and returns notFound("runtime_session_not_found", "Runtime Session is not active in this Broker process") (:2422-2424) when the replacement process has no live context. getExecution (:464-492) reaches this via requireReadySession (:487) for any non-terminal execution — i.e. the in-flight row → 404 → load answers 409 hosted_turn_recovery_required. release (:1015-1033) hits the same null-context path through releasedSession (:1035-1090), which throws unavailable("runtime_reconciliation_required", "Runtime Session is not active in this Broker process") (:1086-1088) unless the binding is LOST with stopped writers and no active executions (:1062-1085) — the settled-but-still-active row falls through to that throw → 503 managed_runtime_cancel_failed. Exactly the two rows in the table above.

3. The unit suite pins the contract the real Broker does not offer. hosted-harness-session.test.ts:4625-4627 stubs acquire to resolve with no-op, :4632 stubs release to resolve, :4827-4836 stubs status/cancel to succeed without any acquire. The passive test reports a parked execution passively and cancels the turn (:4824) then asserts expect(acquireSpy).not.toHaveBeenCalled() (:4842) — i.e. the test requires the passive path to skip acquire — and expect(release).toHaveBeenCalled() (:4873) — i.e. the cancel route releases a lease the Broker never granted. The comment at :4871-4872 states this explicitly. So the suite encodes a Broker contract (status/release succeed unconditionally) that the real RuntimeBrokerService only honours for a session it has adopted, which is precisely what an owner change prevents.

Reachability note — the issue's reachability analysis holds: ManagedAgentStore.insertCancelCommand refuses Workspace Sessions (409 workspace_unavailable) today, so the path is only reachable today by applying the status='CANCELLING' transition directly, then starting owner B. It goes live when Workspace cancel is opened. The cancellation takeover is otherwise contract-complete with no E2E mode, as the #13083 design states.

Fix direction. The candidate (acquire first in the passive branch when Runtime items exist, release-on-failure mirroring the active branch's :358-369 catch) is the right call: it makes the passive path's Broker interactions match the adoption contract that requireReadySession and release actually enforce, and the test assertion flip (acquireSpy called, plus an acquire stub on the three passive tests) corrects the pinned-but-wrong contract. The alternative — Broker-side: answer non-terminal status from the durable record and accept release on a never-adopted session — is broader and overlaps #13040's Broker-side reconciliation loop; the recovery-side fix is the minimal change that closes this path and keeps the Broker's adoption invariant intact.

Evidence re-derived from main at 728c13de (equals #13083 squash on the affected files). Posting as independent confirmation to unblock the candidate patch's review.

added a commit that references this issue on Oct 3, 2026
66a6279
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions