Skip to content

[Workflow Health] daily-fact mempalace fix still missing (tracker auto-expired) + new daily-firewall-report secret-redaction cra #63348

Description

@github-actions

Workflow Health: 2 findings — daily-fact stale-tracker root cause unfixed, new daily-firewall-report crash signature

1. daily-fact.md — mempalace MCP startup race (P2, prior tracker #62868 auto-expired unfixed)

Prior tracker #62868 (filed 2026-09-23) was closed 2026-09-24T06:55:58Z with stateReason: NOT_PLANNED — this is an expires: 1d auto-closure, not a fix. .github/workflows/shared/mcp/mempalace.md is unchanged (no commits since #63312, a docs-only PR). daily-fact has now failed 16/16 of its last scheduled runs back to at least 2026-09-02, with the newest failure at 2026-09-24T14:22:14Z (run 36012299234) — same signature confirmed previously: dial tcp 172.17.0.1:8765: connect: connection refused from the codex MCP-server-check step reaching the backgrounded mempalace server before it's ready.

Suggested fix (unchanged from #62868): add a wait-for-port / readiness probe with retry to the "Start MemPalace MCP Server" step in shared/mcp/mempalace.md before the MCP gateway's health check runs, or surface /tmp/gh-aw/mcp-logs/mempalace/server.log on failure for direct root-cause diagnosis of why the server isn't binding port 8765 in time.

Recommendation: since P2/P1 trackers with expires: 1d keep self-closing before anyone acts (this is now the 2nd occurrence of this specific pattern reported by this workflow), consider exempting workflow-health issues from short expiry windows, or re-filing with a longer expires value.

2. daily-firewall-report.md — new crash signature: secret-redaction stack overflow (P1, previously untracked)

Distinct from the "chart sub-agent model-access error" noted in the previous run's triage (that was a soft/non-fatal issue the agent worked around). This is a hard crash in the post-agent secret-redaction step, confirmed via live job logs on 2 consecutive days:

  • Run 36086384589 (2026-09-25T02:58Z): agent completed successfully (created discussion, printed summary), but the job still failed with:
    ##[error]ERR_VALIDATION: Secret redaction failed: ERR_VALIDATION: Failed to scan directory /tmp/gh-aw: ERR_VALIDATION: Failed to scan directory /tmp/gh-aw/aw-mcp: Maximum call stack size exceeded
  • Run 35947503250 (2026-09-24T02:56Z): identical error, same directory (/tmp/gh-aw/aw-mcp), same Maximum call stack size exceeded.

This crashes the agent job post-execution even though the actual agent work (discussion creation) already succeeded — i.e. a false-negative failure report masking a healthy run. The stack overflow strongly suggests a symlink loop or deeply-recursive directory structure under /tmp/gh-aw/aw-mcp (likely an MCP-gateway working directory specific to this workflow's tool config) that the AWF firewall's secret-redaction directory scanner doesn't guard against (no cycle/depth detection). This is infrastructure-level (AWF firewall action), not this repo's own compiled workflow code — actions/setup/js/redact_secrets.cjs in this repo doesn't reference aw-mcp directly, so the scanning logic lives in the external AWF firewall action.

Suggested fix direction: report upstream to the AWF firewall maintainers to add symlink-loop/max-depth guards to the secret-redaction directory walker, or (as a repo-side mitigation) investigate whether daily-firewall-report.md's tool config creates a self-referential/deeply nested directory under /tmp/gh-aw/aw-mcp that other workflows don't.

Impact: daily-firewall-report is functionally producing correct output (discussion created both days) but reported as a hard failure both days, inflating the workflow's apparent failure rate and triggering unnecessary [aw] ... failed issues (#63332, and yesterday's equivalent).

Compilation status

298/298 workflows have lock files (100%), compile-validate clean — unchanged.

Live re-verification of failing-workflows.json (4 entries)

  • daily-firewall-report: 2 consecutive failures (see finding Add workflow: githubnext/agentics/weekly-research #2 above) — was previously reported as "intermittent"/recovered; now confirmed as a recurring, specific, previously-undiagnosed crash.
  • daily-go-test-parallelizer: fully healthy (5/5 recent runs successful).
  • lint-monster: fully healthy (4/4 most recent runs successful, only 2 stale failures from 2026-09-20/21 in the window).
  • cjs: plain GH Actions CI workflow, out of gh aw scope — unchanged.

Carried-over unfixed items (re-confirmed, see comment on #63098)

Generated by 🏥 Workflow Health Manager - Meta-Orchestrator · copilot · auto · 118.9 AIC · ⌖ 11.1 AIC · ⊞ 12.3K · ◷

  • expires on Sep 25, 2026, 8:46 PM UTC-08:00

Activity

  1. github-actions commented on Sep 26, 2026

    @github-actions
    ContributorAuthor

    This issue was automatically closed because it expired on 2026-09-26T04:46:18.723Z.

    Closed by Workflow

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions