You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
[Workflow Health] daily-fact mempalace fix still missing (tracker auto-expired) + new daily-firewall-report secret-redaction cra #63348
Prior tracker #62868 (filed 2026-09-23) was closed 2026-09-24T06:55:58Z with stateReason: NOT_PLANNED — this is an expires: 1d auto-closure, not a fix. .github/workflows/shared/mcp/mempalace.md is unchanged (no commits since #63312, a docs-only PR). daily-fact has now failed 16/16 of its last scheduled runs back to at least 2026-09-02, with the newest failure at 2026-09-24T14:22:14Z (run 36012299234) — same signature confirmed previously: dial tcp 172.17.0.1:8765: connect: connection refused from the codex MCP-server-check step reaching the backgrounded mempalace server before it's ready.
Suggested fix (unchanged from #62868): add a wait-for-port / readiness probe with retry to the "Start MemPalace MCP Server" step in shared/mcp/mempalace.md before the MCP gateway's health check runs, or surface /tmp/gh-aw/mcp-logs/mempalace/server.log on failure for direct root-cause diagnosis of why the server isn't binding port 8765 in time.
Recommendation: since P2/P1 trackers with expires: 1d keep self-closing before anyone acts (this is now the 2nd occurrence of this specific pattern reported by this workflow), consider exempting workflow-health issues from short expiry windows, or re-filing with a longer expires value.
Distinct from the "chart sub-agent model-access error" noted in the previous run's triage (that was a soft/non-fatal issue the agent worked around). This is a hard crash in the post-agent secret-redaction step, confirmed via live job logs on 2 consecutive days:
Run 36086384589 (2026-09-25T02:58Z): agent completed successfully (created discussion, printed summary), but the job still failed with: ##[error]ERR_VALIDATION: Secret redaction failed: ERR_VALIDATION: Failed to scan directory /tmp/gh-aw: ERR_VALIDATION: Failed to scan directory /tmp/gh-aw/aw-mcp: Maximum call stack size exceeded
Run 35947503250 (2026-09-24T02:56Z): identical error, same directory (/tmp/gh-aw/aw-mcp), same Maximum call stack size exceeded.
This crashes the agent job post-execution even though the actual agent work (discussion creation) already succeeded — i.e. a false-negative failure report masking a healthy run. The stack overflow strongly suggests a symlink loop or deeply-recursive directory structure under /tmp/gh-aw/aw-mcp (likely an MCP-gateway working directory specific to this workflow's tool config) that the AWF firewall's secret-redaction directory scanner doesn't guard against (no cycle/depth detection). This is infrastructure-level (AWF firewall action), not this repo's own compiled workflow code — actions/setup/js/redact_secrets.cjs in this repo doesn't reference aw-mcp directly, so the scanning logic lives in the external AWF firewall action.
Suggested fix direction: report upstream to the AWF firewall maintainers to add symlink-loop/max-depth guards to the secret-redaction directory walker, or (as a repo-side mitigation) investigate whether daily-firewall-report.md's tool config creates a self-referential/deeply nested directory under /tmp/gh-aw/aw-mcp that other workflows don't.
Impact: daily-firewall-report is functionally producing correct output (discussion created both days) but reported as a hard failure both days, inflating the workflow's apparent failure rate and triggering unnecessary [aw] ... failed issues (#63332, and yesterday's equivalent).
Compilation status
298/298 workflows have lock files (100%), compile-validate clean — unchanged.
Live re-verification of failing-workflows.json (4 entries)
daily-firewall-report: 2 consecutive failures (see finding Add workflow: githubnext/agentics/weekly-research #2 above) — was previously reported as "intermittent"/recovered; now confirmed as a recurring, specific, previously-undiagnosed crash.
Workflow Health: 2 findings — daily-fact stale-tracker root cause unfixed, new daily-firewall-report crash signature
1. daily-fact.md — mempalace MCP startup race (P2, prior tracker #62868 auto-expired unfixed)
Prior tracker #62868 (filed 2026-09-23) was closed 2026-09-24T06:55:58Z with
stateReason: NOT_PLANNED— this is anexpires: 1dauto-closure, not a fix..github/workflows/shared/mcp/mempalace.mdis unchanged (no commits since #63312, a docs-only PR). daily-fact has now failed 16/16 of its last scheduled runs back to at least 2026-09-02, with the newest failure at 2026-09-24T14:22:14Z (run 36012299234) — same signature confirmed previously:dial tcp 172.17.0.1:8765: connect: connection refusedfrom the codex MCP-server-check step reaching the backgroundedmempalaceserver before it's ready.Suggested fix (unchanged from #62868): add a wait-for-port / readiness probe with retry to the "Start MemPalace MCP Server" step in
shared/mcp/mempalace.mdbefore the MCP gateway's health check runs, or surface/tmp/gh-aw/mcp-logs/mempalace/server.logon failure for direct root-cause diagnosis of why the server isn't binding port 8765 in time.Recommendation: since P2/P1 trackers with
expires: 1dkeep self-closing before anyone acts (this is now the 2nd occurrence of this specific pattern reported by this workflow), consider exemptingworkflow-healthissues from short expiry windows, or re-filing with a longerexpiresvalue.2. daily-firewall-report.md — new crash signature: secret-redaction stack overflow (P1, previously untracked)
Distinct from the "chart sub-agent model-access error" noted in the previous run's triage (that was a soft/non-fatal issue the agent worked around). This is a hard crash in the post-agent secret-redaction step, confirmed via live job logs on 2 consecutive days:
##[error]ERR_VALIDATION: Secret redaction failed: ERR_VALIDATION: Failed to scan directory /tmp/gh-aw: ERR_VALIDATION: Failed to scan directory /tmp/gh-aw/aw-mcp: Maximum call stack size exceeded/tmp/gh-aw/aw-mcp), sameMaximum call stack size exceeded.This crashes the
agentjob post-execution even though the actual agent work (discussion creation) already succeeded — i.e. a false-negative failure report masking a healthy run. The stack overflow strongly suggests a symlink loop or deeply-recursive directory structure under/tmp/gh-aw/aw-mcp(likely an MCP-gateway working directory specific to this workflow's tool config) that the AWF firewall's secret-redaction directory scanner doesn't guard against (no cycle/depth detection). This is infrastructure-level (AWF firewall action), not this repo's own compiled workflow code —actions/setup/js/redact_secrets.cjsin this repo doesn't referenceaw-mcpdirectly, so the scanning logic lives in the external AWF firewall action.Suggested fix direction: report upstream to the AWF firewall maintainers to add symlink-loop/max-depth guards to the secret-redaction directory walker, or (as a repo-side mitigation) investigate whether
daily-firewall-report.md's tool config creates a self-referential/deeply nested directory under/tmp/gh-aw/aw-mcpthat other workflows don't.Impact: daily-firewall-report is functionally producing correct output (discussion created both days) but reported as a hard failure both days, inflating the workflow's apparent failure rate and triggering unnecessary
[aw] ... failedissues (#63332, and yesterday's equivalent).Compilation status
298/298 workflows have lock files (100%), compile-validate clean — unchanged.
Live re-verification of failing-workflows.json (4 entries)
gh awscope — unchanged.Carried-over unfixed items (re-confirmed, see comment on #63098)
model-provider: github, gpclean.md retiredgpt-5-codexmodel — all 3 reproduced again today via live logs, no fix PR yet despite Copilot assignment attempted 2026-09-24. See [Workflow Health] 3 root-caused workflow failures: avenger npm-symlink, metrics-collector missing model-provider, gpclean retire #63098 for full detail (comment added this run).