Problem statement
In the 24h window ending 2026-10-01 ~23:21 UTC, the agent job's "Redact secrets in logs" step failed on 3 runs across 3 distinct workflows and 2 different agentic engines — in every case after the agent CLI execution step itself had already completed successfully. This is a novel cluster (no existing open issue covers it, searched this session) and meets the fleet-wide monitor's issue-filing threshold (≥2 distinct workflows affected).
Affected workflows / run IDs
- Deep Report (engine=claude) — §36907433969,
Execute Claude Code CLI = success, Redact secrets in logs = failure (step 51)
- Daily Agent of the Day Blog Writer (engine=copilot) — §36876415253,
Execute GitHub Copilot CLI = success, Redact secrets in logs = failure (step 41)
- Daily Ambient Context Optimizer (engine=copilot) — §36916026748,
Execute GitHub Copilot CLI = success, Redact secrets in logs = failure (step 43)
Evidence
run_summary.json's job_details[].steps[] for the agent job on each run above shows the CLI-execution step completing with "conclusion": "success", followed later in the same job by "name": "Redact secrets in logs" with "conclusion": "failure" — the job's overall conclusion is failure even though the model turn itself succeeded and (for at least Daily Agent of the Day Blog Writer) a separate downstream "Handle agent failure" job then also ran.
Raw step stderr for the Redact secrets in logs step itself was not retrievable via mcp__github__get_job_logs in this session — the tool's tail only returned later post-job cleanup output, not the failing step's own console output. A maintainer with direct Actions UI/log access will need to pull the full step log for one of the three run IDs above to get the exact stack trace or exit code.
Probable root cause
Since this spans two different agent engines (Claude and Copilot) rather than clustering on one, the shared secret-redaction script that runs after every engine's CLI step is the more likely fault location, rather than an engine-specific bug. Possible triggers worth checking: an edge case in log content/size for these three specific runs (e.g. unusual characters, very large log volume, or a secret-like pattern that crashes the redaction regex/parser) rather than every run hitting this step.
Blast radius
Fleet-wide at the step level (any workflow reaching "Redact secrets in logs" is exposed), but low-frequency in practice this window: 3 of 203 runs (~1.5%).
Suggested next steps
- Pull the full
Redact secrets in logs step log for §36907433969 (or the other two run IDs) to get the actual error/exit code.
- Compare the agent log content/size for these 3 failing runs against same-day successful runs of the same step, to isolate whether it's a content-triggered edge case vs. a flaky/resource issue (OOM, timeout).
- Add defensive error handling/timeout around the redaction step so a failure there doesn't fail the whole job when the actual agent output was already produced successfully — consider treating redaction failures as a warning with fallback (e.g. skip redaction and flag the artifact) rather than a hard job failure, since the underlying security control (the firewall/network allowlist) already ran clean on all three flagged runs.
Source
Filed from the daily Agent Job Health Monitor report for the 24h window ending 2026-10-01 ~23:21 UTC. Fleet run-weighted agent-job failure rate this window: 11/203 = 5.42% (5 tracked via a separate recurring Copilot-CLI crash pattern, 6 novel including this cluster).
Generated by 🩺 Agent Job Health Monitor · claude · agent · 961.7 AIC · ⌖ 6.92 AIC · ⊞ 6.8K · ◷
Problem statement
In the 24h window ending 2026-10-01 ~23:21 UTC, the
agentjob's "Redact secrets in logs" step failed on 3 runs across 3 distinct workflows and 2 different agentic engines — in every case after the agent CLI execution step itself had already completed successfully. This is a novel cluster (no existing open issue covers it, searched this session) and meets the fleet-wide monitor's issue-filing threshold (≥2 distinct workflows affected).Affected workflows / run IDs
Execute Claude Code CLI= success,Redact secrets in logs= failure (step 51)Execute GitHub Copilot CLI= success,Redact secrets in logs= failure (step 41)Execute GitHub Copilot CLI= success,Redact secrets in logs= failure (step 43)Evidence
run_summary.json'sjob_details[].steps[]for theagentjob on each run above shows the CLI-execution step completing with"conclusion": "success", followed later in the same job by"name": "Redact secrets in logs"with"conclusion": "failure"— the job's overall conclusion isfailureeven though the model turn itself succeeded and (for at least Daily Agent of the Day Blog Writer) a separate downstream "Handle agent failure" job then also ran.Raw step stderr for the
Redact secrets in logsstep itself was not retrievable viamcp__github__get_job_logsin this session — the tool's tail only returned later post-job cleanup output, not the failing step's own console output. A maintainer with direct Actions UI/log access will need to pull the full step log for one of the three run IDs above to get the exact stack trace or exit code.Probable root cause
Since this spans two different agent engines (Claude and Copilot) rather than clustering on one, the shared secret-redaction script that runs after every engine's CLI step is the more likely fault location, rather than an engine-specific bug. Possible triggers worth checking: an edge case in log content/size for these three specific runs (e.g. unusual characters, very large log volume, or a secret-like pattern that crashes the redaction regex/parser) rather than every run hitting this step.
Blast radius
Fleet-wide at the step level (any workflow reaching "Redact secrets in logs" is exposed), but low-frequency in practice this window: 3 of 203 runs (~1.5%).
Suggested next steps
Redact secrets in logsstep log for §36907433969 (or the other two run IDs) to get the actual error/exit code.Source
Filed from the daily Agent Job Health Monitor report for the 24h window ending 2026-10-01 ~23:21 UTC. Fleet run-weighted agent-job failure rate this window: 11/203 = 5.42% (5 tracked via a separate recurring Copilot-CLI crash pattern, 6 novel including this cluster).