You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
[agent-job-health] Agent Job Health Monitor daily report
#67304
Window evaluated: 2026-10-09T15:31:57Z – 2026-10-09T23:21:02Z UTC (~7.8h). The local pre-downloaded log cache did not cover a full 24h window this run — see note below.
Run-weighted fleet failure rate: 3.04% (8/263, confirmed basis). Upper bound if every indeterminate run were an agent failure: 18.25% (48/263).
Workflow-weighted: median 0.0%, mean 13.62% (over the 41 workflows with full job-level data).
Tracked vs. novel: 0 tracked / 8 confirmed-novel agent-job failures (issue search for known signatures returned no accessible matches — see caveat below). Plus 40 runs in an indeterminate cluster (see Novel Failure Clusters).
Trend: Confirmed-basis rate (3.0%) is lower than 2026-10-07 (12.77%), but this is not a like-for-like comparison — today's local sample is much smaller and narrower than prior runs (7.8h vs ~24h), and 40 of today's 61 overall-failed runs could not be attributed to a specific job. Not treating this as a real improvement.
Warning
Two data-quality issues limit confidence in today's numbers: (1) the local log cache covered only ~1/3 of the claimed 24h window, and (2) the GitHub Actions Jobs API returns zero jobs for 40 failed PR-reviewer-fleet runs, making it impossible to confirm whether agent or another job actually failed for those runs.
Failure Rate by Step (confirmed agent-job failures only, n=8)
Failing step
Runs
Share of confirmed agent_failures
Execute Codex CLI
2
25%
Build gh-aw Docker image
2
25%
Execute GitHub Copilot CLI
2
25%
Install Codex CLI
1
12.5%
Warm up model
1
12.5%
Tracked Failures
None. Issue search (search_issues) returned zero accessible results for every query tried against today's failure signatures and the known chronic-offender names (PR Sous Chef, Issue Monster, Contribution Check, Copilot CLI segfault) — all of which appear in today's sample with a 0% failure rate, consistent with them being already-mitigated/tracked. Every other search had all candidate results removed by the integrity policy (resource trust level below "approved"), so absence of a hit here should not be read as confirmation that no tracking issue exists — only that this run could not verify one.
Novel Failure Clusters
Indeterminate PR-reviewer-fleet failures — 40 runs, 8 workflows (AI Moderator, Visual Regression Checker, Matt Pocock Skills Reviewer, Test Quality Sentinel, Impeccable Skills Reviewer, Ponytail Reviewer, Design Decision Gate 🏗️, PR Code Quality Reviewer). Engine: mixed (copilot/codex). Representative run: §37956584779 (AI Moderator, pull_request, conclusion failure). Root cause unknown: both list_workflow_jobs and get_job_logs return total_count: 0 for every one of these runs, and the local log cache never downloaded run_summary.json/aw_info.json for them either — all triggered by pull_request events on Copilot-authored branches (actor Copilot/copilot-swe-agent). Suspected cause: restricted job-log visibility for PR runs opened by the Copilot coding agent, not necessarily the agent job itself failing (could be detection, which is out of scope). Recommended action: investigate why Actions job-level data is unavailable for these runs (token scope? retention? a different workflow structure for Copilot-branch PRs?) before concluding this is an agent-job regression.
Build gh-aw Docker image step failure — 2 runs, 2 workflows (Daily Safe Output Tool Optimizer §37990260860, Agentic Workflow Audit Agent §37992875335), engine: claude. The step failed in ~11s (not a real compile failure — looks like a docker buildx setup/cache error); available log excerpt only shows post-job teardown, not the actual error. Below the 10%/2-workflow issue threshold on its own but worth a follow-up log pull.
Execute Codex CLI / Install Codex CLI step failures — 3 runs total across Daily Go Test Parallelizer (x2) and Daily Cache Strategy Analyzer (x1), single-workflow-dominant, below threshold.
Schedule Heartbeat
No blind spots detected in the checked sample. Because the local log cache covered only ~7.8h instead of 24h, 202 of 241 schedule-triggered workflows had no run in the local sample — this is expected given the narrow window, not evidence of a blind spot by itself. A stratified sample of 32 of those 202 (daily/sub-daily cadence, excluding smoke-* fixtures and weekly+ cadences) was checked directly against the GitHub Actions API (list_workflow_runs, most recent run of any kind): 32/32 had a run within their expected cadence (all within the last ~20h), several with conclusion: failure (already covered under indeterminate/novel clusters above where applicable) but none silent beyond 2× cadence + 1 day. Full exhaustive verification of all 241 scheduled workflows was not completed this run (budget); the remaining ~170 (mostly daily/every 2 days smoke fixtures) were not individually checked.
Workflow (sample)
Last observed run (UTC)
Expected cadence
Gap
ab-testing-advisor
2026-10-09T10:35:00Z
daily
<1 day
cli-version-checker
2026-10-09T05:35:30Z
daily
<1 day
aw-issue-clustering
2026-10-09T07:01:18Z
daily
<1 day
daily-doc-healer
2026-10-08T23:44:22Z
daily
<1 day
(28 more checked, all <1 day gap)
—
—
—
View per-workflow breakdown (workflows with full job-level data, n=41)
Workflow
Runs
Agent failures
Rate
Daily Go Test Parallelizer
4
2
50%
Agentic Workflow Audit Agent
1
1
100%
Daily BYOK Ollama Test
1
1
100%
Daily Cache Strategy Analyzer
1
1
100%
Daily Reliability Review
1
1
100%
Daily Safe Output Tool Optimizer
1
1
100%
Matt Pocock Skills Reviewer
12
1
8.3%
PR Sous Chef
16
0
0%
PR Code Quality Reviewer
14
0
0%
Test Quality Sentinel
14
0
0%
Ponytail Reviewer
14
0
0%
Design Decision Gate 🏗️
13
0
0%
Impeccable Skills Reviewer
12
0
0%
Issue Monster
10
0
0%
Auto-Triage Issues
8
0
0%
(27 more workflows, 0 failures each)
—
0
0%
Recommendations
Investigate the Actions job-visibility gap for pull_request-triggered runs on Copilot-authored branches (40 runs this window alone) — this blocks accurate agent-job attribution for a meaningful share of the PR-reviewer fleet and should be fixed before the next report, not worked around again.
Pull full logs for the "Build gh-aw Docker image" failures (Daily Safe Output Tool Optimizer, Agentic Workflow Audit Agent) — the ~11s failure duration suggests a buildx/cache setup issue, not a real compile break; worth a quick fix before it spreads to more workflows.
No action needed on the schedule front — sampled heartbeat check found 32/32 healthy; recommend a future run budget more log-download capacity so the local cache actually spans 24h, which would make both the fleet rate and the heartbeat check load-bearing without manual API backfill.
Re-run issue-tracking lookups with elevated integrity scope if available — this run's search_issues calls were filtered for nearly every query, so the tracked/novel split above should be treated as a lower bound on "tracked."
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
Summary
Warning
Two data-quality issues limit confidence in today's numbers: (1) the local log cache covered only ~1/3 of the claimed 24h window, and (2) the GitHub Actions Jobs API returns zero jobs for 40 failed PR-reviewer-fleet runs, making it impossible to confirm whether
agentor another job actually failed for those runs.Failure Rate by Step (confirmed agent-job failures only, n=8)
Tracked Failures
None. Issue search (
search_issues) returned zero accessible results for every query tried against today's failure signatures and the known chronic-offender names (PR Sous Chef, Issue Monster, Contribution Check, Copilot CLI segfault) — all of which appear in today's sample with a 0% failure rate, consistent with them being already-mitigated/tracked. Every other search had all candidate results removed by the integrity policy (resource trust level below "approved"), so absence of a hit here should not be read as confirmation that no tracking issue exists — only that this run could not verify one.Novel Failure Clusters
pull_request, conclusionfailure). Root cause unknown: bothlist_workflow_jobsandget_job_logsreturntotal_count: 0for every one of these runs, and the local log cache never downloadedrun_summary.json/aw_info.jsonfor them either — all triggered bypull_requestevents on Copilot-authored branches (actorCopilot/copilot-swe-agent). Suspected cause: restricted job-log visibility for PR runs opened by the Copilot coding agent, not necessarily theagentjob itself failing (could bedetection, which is out of scope). Recommended action: investigate why Actions job-level data is unavailable for these runs (token scope? retention? a different workflow structure for Copilot-branch PRs?) before concluding this is an agent-job regression.docker buildxsetup/cache error); available log excerpt only shows post-job teardown, not the actual error. Below the 10%/2-workflow issue threshold on its own but worth a follow-up log pull.Schedule Heartbeat
No blind spots detected in the checked sample. Because the local log cache covered only ~7.8h instead of 24h, 202 of 241 schedule-triggered workflows had no run in the local sample — this is expected given the narrow window, not evidence of a blind spot by itself. A stratified sample of 32 of those 202 (daily/sub-daily cadence, excluding
smoke-*fixtures and weekly+ cadences) was checked directly against the GitHub Actions API (list_workflow_runs, most recent run of any kind): 32/32 had a run within their expected cadence (all within the last ~20h), several withconclusion: failure(already covered under indeterminate/novel clusters above where applicable) but none silent beyond 2× cadence + 1 day. Full exhaustive verification of all 241 scheduled workflows was not completed this run (budget); the remaining ~170 (mostlydaily/every 2 dayssmoke fixtures) were not individually checked.View per-workflow breakdown (workflows with full job-level data, n=41)
Recommendations
pull_request-triggered runs on Copilot-authored branches (40 runs this window alone) — this blocks accurate agent-job attribution for a meaningful share of the PR-reviewer fleet and should be fixed before the next report, not worked around again.search_issuescalls were filtered for nearly every query, so the tracked/novel split above should be treated as a lower bound on "tracked."All reactions