Skip to content

[AW Top 10] 07 Investigate synchronized hangs in PR review workflows #66819

Description

@github-actions

Priority 7/10 | 7 source issues | Impact 3/5 | Confidence 3/5 | Effort 3/5

One assignment, one coherent fix

Several PR review persona workflows start at identical times and then time out or hang until the job cap.

Implementation scope

Identify the shared trigger or dependency, add timeout guards and missing-tool preflight checks, and fix the root cause so the reviewers finish within their budget.

Done when

  • The persona workflows complete within their configured timeouts on a pull request.
  • A hang produces a timeout with diagnostic output rather than running to the job cap.

Why now

The timeouts recur on pull request branches and the synchronized start times suggest a common cause; confidence is moderate because individual runs were not reproduced.

AW source issues and corroborating reports

#53619 #66238 #66241 #66274 #66297 #67292 #67299

No corroborating AW discussion; evidence comes from the source issues.

Unchanged AW sources close only after this summary is completed. Newer source activity and not-planned retirement do not trigger source closure. Assigned summaries are frozen; unassign to allow reclustering.

Generated by AW Essential Issue Clustering · copilot · auto · 28.2 AIC · ⌖ 0.588 AIC · ⊞ 8.9K · ◷

Activity

  1. changed the title [-][AW Top 10] 05 Investigate synchronized hangs in PR review workflows[/-] [+][AW Top 10] 07 Investigate synchronized hangs in PR review workflows[/+] on Oct 9, 2026
  2. pelikhan commented on Oct 10, 2026

    @pelikhan
    Collaborator

    RCA — Synchronized hangs/timeouts in PR review persona workflows

    Status: Outstanding, actively reproducing (not fixed)

    Scope note: Of the 7 clustered source issues, two patterns were found and they are not the same root cause:

    1. Primary, currently-reproducing pattern — Impeccable Skills Reviewer + Matt Pocock Skills Reviewer ([aw] Failed jobs: Impeccable Skills Reviewer #66238, [aw] Failed jobs: Matt Pocock Skills Reviewer #66241, [aw] Failed jobs: Impeccable Skills Reviewer #66274, [aw] Failed jobs: Impeccable Skills Reviewer #66297, [aw] Matt Pocock Skills Reviewer timed out #67292, [aw] Impeccable Skills Reviewer timed out #67299).
    2. Stale/unrelated — Design Decision Gate ([aw-failures] [P1] Design Decision Gate hangs silently mid-tool-use in Claude Code CLI step #53619). Its own follow-up comment already retracted the "silent hang" framing (it was a 30/30 LLM-invocation-cap failure). A fresh Design Decision Gate failure I checked (run 38044878675, 2026-10-10) had its agent job succeed in 3 min; the failure was in safe_outputs, unrelated to hangs. Recommend dropping [aw-failures] [P1] Design Decision Gate hangs silently mid-tool-use in Claude Code CLI step #53619 from this cluster.

    Root cause (confidence 8/10)

    .github/workflows/impeccable-skills-reviewer.md and .github/workflows/mattpocock-skills-reviewer.md both hardcode:

    timeout-minutes: 15

    on the engine execution step (compiled as the Execute GitHub Copilot CLI step, e.g. impeccable-skills-reviewer.lock.yml:1092 / mattpocock-skills-reviewer.lock.yml:1195). This is separate from the job-level timeout-minutes: 60.

    Verified via gh run view <id> --log on runs 37502435413, 37523691712, 37535185512:

    • Each shows [detect-agent-errors] Detected step timeout: the engine execution step reached its 15-minute timeout-minutes budget and was terminated and categories=["agentic_engine_timeout"] — a hard kill at exactly 15 minutes, not an indefinite hang. numTurns≈87 at kill time, i.e. actively working (grep/skill lookups/diff review), not blocked on one call.
    • gh run list shows ~30–40% failure rate for both workflows, not 100% — consistent with timeout being diff/PR-size dependent.
    • Synchronization explained: both workflows trigger on the same pull_request: ready_for_review event and review the same diff, so a sufficiently large/complex PR times out both within seconds of each other — this produced the "synchronized start times" signal, but it's shared triggering + a too-tight shared budget, not a shared external dependency.
    • timeout-minutes: 15 has been unchanged since these files were created (git log -p), never recalibrated as skill-driven review flow (prefetch caching, deterministic skill selection, etc.) grew more turn-intensive over time.

    Secondary, narrower symptom (#67292/#67299 only): those runs show fatal: unable to access 'https://github.com/github/gh-aw.git/': server certificate verification failed from git log/git show run inside the sandboxed CLI — expected firewall behavior (no direct git remote access from inside the sandbox), but the agent burned turns retrying instead of fast-failing, helping it hit the 15-min ceiling sooner. Contributing factor for those two runs only, not the general root cause.

    Minimal fix

    1. Raise timeout-minutes on the agent engine step in both .md files from 15 to 25–30 (job-level timeout-minutes: 60 stays as outer bound); gh aw compile/make recompile to regenerate the .lock.yml files.
    2. Confirm the run-failure message template surfaces "timed out" explicitly (detect-agent-errors already emits agentic_engine_timeout) so future occurrences are self-diagnosing without a full log pull.
    3. For [aw] Matt Pocock Skills Reviewer timed out #67292/[aw] Impeccable Skills Reviewer timed out #67299: no fix needed in the reviewed PRs; consider a short note in shared/pr-diff-data-fetch.md telling the agent not to run raw git log/git show against the remote (local checkout only), to stop wasted turns on failing network git calls.
    4. Drop [aw-failures] [P1] Design Decision Gate hangs silently mid-tool-use in Claude Code CLI step #53619 from this essential cluster when triaging source issues — stale/already explained.

    Effort

    Low (2/10) — two one-line frontmatter edits + recompile; no Go/runtime code changes.

    Risks

    Raising the timeout increases worst-case latency/AI-credit spend per run for large PRs, bounded by the existing job-level 60-minute cap. Does not address why skill-driven reviews need this many turns — if turn counts keep growing this will need revisiting (e.g. diff-size capping, as previously proposed-but-superseded in #53619).

    Validation

    1. Bump timeout-minutes: 15 → 25 in both .md files, make recompile, confirm only the Execute GitHub Copilot CLI step timeout changes (job-level 60 unchanged).
    2. Re-trigger both workflows on a PR that previously hit [aw] Failed jobs: Impeccable Skills Reviewer #66274/[aw] Failed jobs: Impeccable Skills Reviewer #66297 — confirm agent job concludes without agentic_engine_timeout.
    3. gh run view <new_run_id> --log | grep agentic_engine_timeout → no match; numTurns completes under budget.
    4. Monitor next ~10 runs of each workflow for failure-rate drop from the observed ~30–40%.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    agentic-workflowsautomationaw-essentialEssential AW-generated issue clusters: assign one to resolve related findingscookieIssue Monster Loves Cookies!

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions