Novel failure cluster: "Redact secrets in logs" step fails across 6 unrelated workflows
Found by: Agent Job Health Monitor, window ~2026-10-02 08:24–23:34 UTC (~15h sample)
Error signature: the agent job's "Redact secrets in logs" step reports conclusion: failure while the preceding Execute <engine> CLI step in the same job reports success. This is not a single workflow's prompt/logic failing — it's a shared post-processing step failing across engines and workflows.
Affected workflows / representative runs (6 runs, 6 distinct workflows, all in this one ~15h sample):
Measured fleet context: in this sample, total_runs=212 completed runs across distinct_workflows=85, with agent_failures=25 (11.8% run-weighted fleet failure rate). This cluster accounts for 6 of the 25 agent-job failures (24%).
Tracked/novel split: search_issues for this signature (and for "Avenger"/"Codex CLI" and "Copilot CLI segfault" queries) returned no directly-readable open matches in this session — some candidate results were filtered by integrity policy before they could be read to confirm status. This cluster is reported as novel on that basis; please dedupe against any existing issue if one turns out to already cover it.
Suspected root cause: a bug or resource/size limit in the shared log-redaction step, independent of which agentic engine (Copilot/Codex/Claude/Pi/etc.) or workflow produced the log — worth pulling the actual step log from one of the runs above to find the specific error.
Recommended action: investigate the redaction step implementation for an edge case (e.g. log size, encoding, or a specific secret-pattern regex) that's intermittently triggered across otherwise-unrelated workflows, since a fix here would clear this entire cluster at once rather than requiring six separate workflow-level fixes.
Also flagged in the accompanying report but not separately filed here (doesn't meet the ≥2-workflow / ≥10%-of-runs bar for auto-filing): the Avenger workflow failed its agent job's "Execute Codex CLI" step on 6 of 6 observed runs (100%) in the same window — e.g. https://github.com/github/gh-aw/actions/runs/37013841183. Single-workflow, but worth a look given the 100% reproduction rate.
Generated by 🩺 Agent Job Health Monitor · claude · agent · 807.2 AIC · ⌖ 6.81 AIC · ⊞ 6.8K · ◷
Novel failure cluster: "Redact secrets in logs" step fails across 6 unrelated workflows
Found by: Agent Job Health Monitor, window ~2026-10-02 08:24–23:34 UTC (~15h sample)
Error signature: the
agentjob's "Redact secrets in logs" step reportsconclusion: failurewhile the precedingExecute <engine> CLIstep in the same job reportssuccess. This is not a single workflow's prompt/logic failing — it's a shared post-processing step failing across engines and workflows.Affected workflows / representative runs (6 runs, 6 distinct workflows, all in this one ~15h sample):
Measured fleet context: in this sample,
total_runs=212completed runs acrossdistinct_workflows=85, withagent_failures=25(11.8% run-weighted fleet failure rate). This cluster accounts for 6 of the 25 agent-job failures (24%).Tracked/novel split:
search_issuesfor this signature (and for "Avenger"/"Codex CLI" and "Copilot CLI segfault" queries) returned no directly-readable open matches in this session — some candidate results were filtered by integrity policy before they could be read to confirm status. This cluster is reported as novel on that basis; please dedupe against any existing issue if one turns out to already cover it.Suspected root cause: a bug or resource/size limit in the shared log-redaction step, independent of which agentic engine (Copilot/Codex/Claude/Pi/etc.) or workflow produced the log — worth pulling the actual step log from one of the runs above to find the specific error.
Recommended action: investigate the redaction step implementation for an edge case (e.g. log size, encoding, or a specific secret-pattern regex) that's intermittently triggered across otherwise-unrelated workflows, since a fix here would clear this entire cluster at once rather than requiring six separate workflow-level fixes.
Also flagged in the accompanying report but not separately filed here (doesn't meet the ≥2-workflow / ≥10%-of-runs bar for auto-filing): the Avenger workflow failed its
agentjob's "Execute Codex CLI" step on 6 of 6 observed runs (100%) in the same window — e.g. https://github.com/github/gh-aw/actions/runs/37013841183. Single-workflow, but worth a look given the 100% reproduction rate.