Skip to content

Add root-cause diagnostics to generic failure issues - #67471

Merged
pelikhan merged 7 commits into
mainfrom
copilot/aw-top-10-add-root-cause-detail
Oct 10, 2026
Merged

pelikhan merged 7 commits into
mainfrom
copilot/aw-top-10-add-root-cause-detail

Conversation

Copilot AI commented Oct 10, 2026 •

Copy link
Copy Markdown
Contributor

Generic failure issues often lack enough detail to triage. Missing Actions job data also leaves failures unattributed without explaining why.

  • Step diagnostics: Include the failing Agent step and its log tail, capped at 50 lines and 8,000 characters; exclude subsequent-step output.
  • Attribution fallback: Explain missing or inaccessible job data and report the Agent conclusion from workflow dependency results. Preserve the step name when logs are unavailable.
  • Safe rendering: Retain secret masks through normalization before truncation and fence excerpts against Markdown injection.

Copilot AI and others added 2 commits October 10, 2026 16:38
…sues

Co-authored-by: pelikhan <4175913+pelikhan@users.noreply.github.com>
Co-authored-by: pelikhan <4175913+pelikhan@users.noreply.github.com>
Copilot AI changed the title [WIP] Add root-cause detail to generic failure issues Add root-cause diagnostics to generic failure issues Oct 10, 2026
Copilot AI requested a review from pelikhan October 10, 2026 16:42
@pelikhan
pelikhan marked this pull request as ready for review October 10, 2026 17:56
Copilot AI balanced review requested due to automatic review settings October 10, 2026 17:56
@github-actions

github-actions Bot commented Oct 10, 2026 •

Copy link
Copy Markdown
Contributor

🧠 Matt Pocock Skills Reviewer has completed the skills-based review. ✅

🧠 Reviewed using Matt Pocock's skills by Matt Pocock Skills Reviewer

@github-actions

github-actions Bot commented Oct 10, 2026 •

Copy link
Copy Markdown
Contributor

✅ Test Quality Sentinel completed test quality analysis.

Test Quality Sentinel skipped because pre-fetch PR data was unavailable: unable to fetch test file diff

🧪 Test quality analysis by Test Quality Sentinel

@github-actions

github-actions Bot commented Oct 10, 2026 •

Copy link
Copy Markdown
Contributor

✅ Ponytail Reviewer completed successfully!

Lean already. Ship.

Generated by Ponytail Reviewer for #67471

@github-actions

Copy link
Copy Markdown
Contributor

🔎 PR Code Quality Reviewer is reviewing code quality for this pull request...

@github-actions

github-actions Bot commented Oct 10, 2026 •

Copy link
Copy Markdown
Contributor

✅ Design Decision Gate 🏗️ completed the design decision gate check. See the comment below for the result and any generated ADR draft.

No ADR enforcement needed for PR #67471: the PR does not carry the "implementation" label and has 0 new lines in default business logic directories (threshold 100, no custom .design-gate.yml). Evidence: /tmp/gh-aw/agent/adr-prefetch-summary.json (has_implementation_label=false, default_business_additions=0, requires_adr_by_default_volume=false, file_count=2).

🏗️ ADR gate enforced by Design Decision Gate 🏗️

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

Diagnostics can contradict captured logs and falsely report unavailable attribution for successful infrastructure-error runs.

1 open finding
What changed in this PR

Adds root-cause diagnostics to generic Agent failure issues.

Changes:

  • Captures bounded, redacted failing-step log excerpts.
  • Adds attribution fallbacks when job data is unavailable.
  • Expands tests for truncation, masking, API failures, and rendering.
File Description
actions/​setup/​js/​handle_agent_failure.cjs Collects and renders failure diagnostics.
actions/​setup/​js/​handle_agent_failure.test.cjs Tests diagnostic collection and issue output.

🧠 Review effort: Balanced


💡 Add a code-review agent skill for context-aware, tailored reviews. Learn more in the docs.

Comment thread actions/setup/js/handle_agent_failure.cjs Outdated

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Skills-Based Review 🧠

Applied /codebase-design and /tdd. This is a well-scoped, well-tested change — approving with one minor efficiency note.

📋 Key Themes & Highlights

Key Themes

  • Minor perf nit: getFailedAgentDiagnostics redacts the log excerpt twice in a row (see inline comment) — not a bug, just avoidable duplicate work on a hot path.

Positive Highlights

  • ✅ /tdd: Excellent test coverage — boundary conditions (failure-completion-second overlap), binary log formats (ArrayBuffer/Uint8Array), secret redaction across normalization layers, fence-injection prevention, and the dependency-results fallback are all exercised with clear, descriptive test names.
  • ✅ /codebase-design: The refactor from a bare string return to a structured {failingStep, logExcerpt, attributionUnavailable, logUnavailable} result is a clean, backward-compatible extension — call sites use spread (...failureDiagnostics) so buildFailureDiagnosticsContext didn't need restructuring.
  • ✅ Correctly reuses the existing redactAndBoundDiagnostics / renderErrorDetails / collectAddMaskedValues infrastructure instead of inventing new sanitization logic, keeping the module's security invariants consistent.
  • ✅ Secret masks declared after the failing step's log window are still picked up via collectAddMaskedValues(log) scanning the whole log, closing a real gap.

I reviewed /tmp/gh-aw/agent/pr-diff.patch directly (421 lines, 2 files) rather than invoking pr-triage, since the diff was small and self-contained enough for direct review.

🧠 Reviewed using Matt Pocock's skills by Matt Pocock Skills Reviewer · copilot · sonnet50 · 78.6 AIC · ⌖ 14.5 AIC · ⊞ 10.3K
Comment /matt to run again

Comment thread actions/setup/js/handle_agent_failure.cjs Outdated
@pelikhan

Copy link
Copy Markdown
Collaborator

@copilot must be privacy preserving and agent generated strings must be validated Z

Copilot AI and others added 2 commits October 10, 2026 18:37
Co-authored-by: pelikhan <4175913+pelikhan@users.noreply.github.com>
Co-authored-by: pelikhan <4175913+pelikhan@users.noreply.github.com>

Copilot AI commented Oct 10, 2026

Copy link
Copy Markdown
Contributor Author

@copilot must be privacy preserving and agent generated strings must be validated Z

Addressed in bd6c81b: job conclusions are allowlisted, step names are safely code-rendered, and ZIP log excerpts are extracted in a private temporary directory before masking and truncation. The 490 targeted tests pass.

pelikhan and others added 2 commits October 10, 2026 13:21
Use Actions job metadata and validated driver exit codes only. Omit report_incomplete prose and log excerpts from generic failure issues and comments.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Group job, step, and driver metadata in one diagnostics section. Collapse attribution explanations, shorten incompletion guidance, and retain the explicit agent-text exclusion.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[AW Top 10] 03 Add root-cause detail to generic failure issues

3 participants