Repository navigation
[AW Top 10] 07 Investigate synchronized hangs in PR review workflows #66819
Description
Activity
- addedcookieIssue Monster Loves Cookies!Issue Monster Loves Cookies!aw-essentialEssential AW-generated issue clusters: assign one to resolve related findingsEssential AW-generated issue clusters: assign one to resolve related findings
on Oct 8, 2026 - changed the title
[-][AW Top 10] 05 Investigate synchronized hangs in PR review workflows[/-][+][AW Top 10] 07 Investigate synchronized hangs in PR review workflows[/+]on Oct 9, 2026 RCA — Synchronized hangs/timeouts in PR review persona workflows
Status: Outstanding, actively reproducing (not fixed)
Scope note: Of the 7 clustered source issues, two patterns were found and they are not the same root cause:
- Primary, currently-reproducing pattern — Impeccable Skills Reviewer + Matt Pocock Skills Reviewer ([aw] Failed jobs: Impeccable Skills Reviewer #66238, [aw] Failed jobs: Matt Pocock Skills Reviewer #66241, [aw] Failed jobs: Impeccable Skills Reviewer #66274, [aw] Failed jobs: Impeccable Skills Reviewer #66297, [aw] Matt Pocock Skills Reviewer timed out #67292, [aw] Impeccable Skills Reviewer timed out #67299).
- Stale/unrelated — Design Decision Gate ([aw-failures] [P1] Design Decision Gate hangs silently mid-tool-use in Claude Code CLI step #53619). Its own follow-up comment already retracted the "silent hang" framing (it was a
30/30LLM-invocation-cap failure). A fresh Design Decision Gate failure I checked (run 38044878675, 2026-10-10) had itsagentjob succeed in 3 min; the failure was insafe_outputs, unrelated to hangs. Recommend dropping [aw-failures] [P1] Design Decision Gate hangs silently mid-tool-use in Claude Code CLI step #53619 from this cluster.
Root cause (confidence 8/10)
.github/workflows/impeccable-skills-reviewer.mdand.github/workflows/mattpocock-skills-reviewer.mdboth hardcode:timeout-minutes: 15
on the engine execution step (compiled as the
Execute GitHub Copilot CLIstep, e.g.impeccable-skills-reviewer.lock.yml:1092/mattpocock-skills-reviewer.lock.yml:1195). This is separate from the job-leveltimeout-minutes: 60.Verified via
gh run view <id> --logon runs 37502435413, 37523691712, 37535185512:- Each shows
[detect-agent-errors] Detected step timeout: the engine execution step reached its 15-minute timeout-minutes budget and was terminatedandcategories=["agentic_engine_timeout"]— a hard kill at exactly 15 minutes, not an indefinite hang. numTurns≈87 at kill time, i.e. actively working (grep/skill lookups/diff review), not blocked on one call. gh run listshows ~30–40% failure rate for both workflows, not 100% — consistent with timeout being diff/PR-size dependent.- Synchronization explained: both workflows trigger on the same
pull_request: ready_for_reviewevent and review the same diff, so a sufficiently large/complex PR times out both within seconds of each other — this produced the "synchronized start times" signal, but it's shared triggering + a too-tight shared budget, not a shared external dependency. timeout-minutes: 15has been unchanged since these files were created (git log -p), never recalibrated as skill-driven review flow (prefetch caching, deterministic skill selection, etc.) grew more turn-intensive over time.
Secondary, narrower symptom (#67292/#67299 only): those runs show
fatal: unable to access 'https://github.com/github/gh-aw.git/': server certificate verification failedfromgit log/git showrun inside the sandboxed CLI — expected firewall behavior (no direct git remote access from inside the sandbox), but the agent burned turns retrying instead of fast-failing, helping it hit the 15-min ceiling sooner. Contributing factor for those two runs only, not the general root cause.Minimal fix
- Raise
timeout-minuteson the agent engine step in both.mdfiles from15to25–30(job-leveltimeout-minutes: 60stays as outer bound);gh aw compile/make recompileto regenerate the.lock.ymlfiles. - Confirm the
run-failuremessage template surfaces "timed out" explicitly (detect-agent-errors already emitsagentic_engine_timeout) so future occurrences are self-diagnosing without a full log pull. - For [aw] Matt Pocock Skills Reviewer timed out #67292/[aw] Impeccable Skills Reviewer timed out #67299: no fix needed in the reviewed PRs; consider a short note in
shared/pr-diff-data-fetch.mdtelling the agent not to run rawgit log/git showagainst the remote (local checkout only), to stop wasted turns on failing network git calls. - Drop [aw-failures] [P1] Design Decision Gate hangs silently mid-tool-use in Claude Code CLI step #53619 from this essential cluster when triaging source issues — stale/already explained.
Effort
Low (2/10) — two one-line frontmatter edits + recompile; no Go/runtime code changes.
Risks
Raising the timeout increases worst-case latency/AI-credit spend per run for large PRs, bounded by the existing job-level 60-minute cap. Does not address why skill-driven reviews need this many turns — if turn counts keep growing this will need revisiting (e.g. diff-size capping, as previously proposed-but-superseded in #53619).
Validation
- Bump
timeout-minutes: 15→25in both.mdfiles,make recompile, confirm only theExecute GitHub Copilot CLIstep timeout changes (job-level 60 unchanged). - Re-trigger both workflows on a PR that previously hit [aw] Failed jobs: Impeccable Skills Reviewer #66274/[aw] Failed jobs: Impeccable Skills Reviewer #66297 — confirm
agentjob concludes withoutagentic_engine_timeout. gh run view <new_run_id> --log | grep agentic_engine_timeout→ no match; numTurns completes under budget.- Monitor next ~10 runs of each workflow for failure-rate drop from the observed ~30–40%.
- added a commit that references this issue
on Oct 10, 2026
Priority 7/10 | 7 source issues | Impact 3/5 | Confidence 3/5 | Effort 3/5
One assignment, one coherent fix
Several PR review persona workflows start at identical times and then time out or hang until the job cap.
Implementation scope
Identify the shared trigger or dependency, add timeout guards and missing-tool preflight checks, and fix the root cause so the reviewers finish within their budget.
Done when
Why now
The timeouts recur on pull request branches and the synchronized start times suggest a common cause; confidence is moderate because individual runs were not reproduced.
AW source issues and corroborating reports
#53619 #66238 #66241 #66274 #66297 #67292 #67299
No corroborating AW discussion; evidence comes from the source issues.
Unchanged AW sources close only after this summary is completed. Newer source activity and not-planned retirement do not trigger source closure. Assigned summaries are frozen; unassign to allow reclustering.