Skip to content

Add model and engine misconfiguration diagnostics to agentic-workflows - #67488

Merged
pelikhan merged 13 commits into
mainfrom
copilot/add-checklist-for-misconfiguration
Oct 10, 2026
Merged

pelikhan merged 13 commits into
mainfrom
copilot/add-checklist-for-misconfiguration

Conversation

Copilot AI commented Oct 10, 2026 •

Copy link
Copy Markdown
Contributor

Model/endpoint failures led the skill to recommend older models before diagnosing outdated gh-aw compilers and wire-API mismatches. Misplaced CLI and AWF version pins also lacked actionable guidance.

  • Routing and checklists: Route symptoms directly to aligned guidance in the full debug guide and root debug.md.
  • Diagnosis and remedies: Explain prerelease installation, alias resolution, wire-API precedence, session-wide endpoint constraints, and relevant artifacts. Prioritize upgrading and recompiling, then compatible model selection; reserve downgrades for last resort.
  • Version fields: Map Copilot CLI install failures to engine.version and AWF failures to sandbox.agent.version; explicitly reject engine.copilot.version.
  • Regression coverage: Add redacted customer-case fixtures and contract tests guarding remedy order, version-field mapping, artifact guidance, and routing.

Copilot AI and others added 2 commits October 10, 2026 18:33
Co-authored-by: SivaKesava1 <11771739+SivaKesava1@users.noreply.github.com>
Co-authored-by: SivaKesava1 <11771739+SivaKesava1@users.noreply.github.com>
Copilot AI changed the title [WIP] Add checklist for diagnosing model and engine misconfiguration Add model and engine misconfiguration diagnostics to agentic-workflows Oct 10, 2026
Copilot AI requested a review from SivaKesava1 October 10, 2026 18:36
@SivaKesava1
SivaKesava1 marked this pull request as ready for review October 10, 2026 18:36
Copilot AI balanced review requested due to automatic review settings October 10, 2026 18:36
@github-actions

Copy link
Copy Markdown
Contributor

🧠 Matt Pocock Skills Reviewer is reviewing this pull request using Matt Pocock's engineering skills...

@github-actions

Copy link
Copy Markdown
Contributor

🔬 Test Quality Sentinel is analyzing test quality on this pull request...

@github-actions

Copy link
Copy Markdown
Contributor

✂️ Ponytail Reviewer has started processing this pull request

@github-actions

Copy link
Copy Markdown
Contributor

🔎 PR Code Quality Reviewer is reviewing code quality for this pull request...

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

The new regression suite is not registered with any required Make or CI test target.

1 open finding
What changed in this PR

Adds targeted diagnostics for model, endpoint, and version-pin misconfiguration in agentic workflows.

Changes:

  • Routes relevant failures to a new debugging checklist.
  • Documents compiler, wire-API, version-field, and artifact diagnosis.
  • Adds documentation contract fixtures and tests.
File Description
debug.md Adds condensed public troubleshooting guidance.
.github/​aw/​debug-agentic-workflow.md Adds the comprehensive diagnostic checklist.
.github/​skills/​agentic-workflows/​SKILL.md Routes matching symptoms to the checklist.
.github/​skills/​agentic-workflows/​tests/​test_guidance.py Adds documentation contract tests.
.github/​skills/​agentic-workflows/​tests/​fixtures.json Adds synthetic regression scenarios.

🧠 Review effort: Balanced


💡 Add a code-review agent skill for context-aware, tailored reviews. Learn more in the docs.

Comment thread .github/skills/agentic-workflows/tests/test_guidance.py
@SivaKesava1

Copy link
Copy Markdown
Collaborator

@copilot Please merge main. The red agentic workflow checks on this PR came from a safe outputs regression on main (the safe outputs MCP server could not load memory_schema_contract.cjs), which #67506 fixed. Please keep this PR to the #67486 scope and don't fix unrelated failures here. We'll review the checklist content once the checks are green.

…-for-misconfiguration

Co-authored-by: SivaKesava1 <11771739+SivaKesava1@users.noreply.github.com>

Copilot AI commented Oct 10, 2026

Copy link
Copy Markdown
Contributor Author

@copilot Please merge main. The red agentic workflow checks on this PR came from a safe outputs regression on main (the safe outputs MCP server could not load memory_schema_contract.cjs), which ...

Merged main (including #67506) into this PR in merge commit 53e9b864a8, keeping the changes scoped to #67486.

@gh-aw-bot

Copy link
Copy Markdown
Collaborator

@copilot address the following outstanding work in one pass:

  1. Update this branch with the latest main using make merge-main, resolving any conflicts and preserving the intended changes.
  2. Review (.github/skills/agentic-workflows/tests/test_guidance.py:92): These contract tests are only executable manually: no Make target or CI workflow discovers this directory, and test-all covers Go, JavaScript, WASM, and shell tests only. As a result, guidance can regress without this suite failing, so the promised regression guard is ineffective until this command is wired into a required test target or workflow. - Add model and engine misconfiguration diagnostics to agentic-workflows #67488 (comment)

Push the necessary fixes, reply to each listed review thread and resolve it when addressed. Ignore feedback already answered or resolved. Use the pr-finisher skill and stop when only human review or CI remains; do not trigger CI.

Sous-chef head: 53e9b86
Sous-chef work: 260832280d089284589673418b5b6ec56f78f5eaa47828823e7adffd46fa4fb4
Sous-chef state: 14b8584b9c5fdffacbace89f6295c20c6b8c8a4d3c5dcaff6bfce01ff66da8c1

Generated by 👨‍🍳 PR Sous Chef · pi · haiku45 · 3.99 AIC · ⌖ 7.65 AIC · ⊞ 1K · ◷
Comment /souschef to run again

Co-authored-by: gh-aw-bot <259018956+gh-aw-bot@users.noreply.github.com>

Copilot AI commented Oct 10, 2026

Copy link
Copy Markdown
Contributor Author

@copilot address the following outstanding work in one pass:

  1. Update this branch with the latest main using make merge-main, resolving any conflicts and preserving the intended changes....

make merge-main completed; the branch already contained the latest main. Added make test-guidance and wired it into test-all; validation passed and the change is in 58752b2. I replied to the requested review thread, but its resolution could not be performed: the available GitHub tools do not expose thread resolution, and gh has no GH_TOKEN in this environment. CI was not triggered.

Copilot AI requested a review from gh-aw-bot October 10, 2026 21:05
@SivaKesava1

Copy link
Copy Markdown
Collaborator

@copilot Content review of the checklist (in addition to gh-aw-bot's point about running the guidance tests in CI):

  1. The wire API order in step 2 describes main before Select Copilot wire APIs from AWF metadata and surface model mismatches #67487. Select Copilot wire APIs from AWF metadata and surface model mismatches #67487 (open) makes AWF's per-model endpoint metadata (/reflect supported_endpoints) take precedence and fail fast on a mismatch at startup (Model endpoint mismatch: … Pin a compatible model, remove the COPILOT_PROVIDER_WIRE_API override, or upgrade gh-aw.). Either describe the order in a way that stays true after that lands, or add the AWF metadata step and the startup mismatch message.
  2. "One wire API for the whole session, including sub-agents" holds for the default CLI mode only. With Fix Copilot SDK model routing and cross-family sub-agents #67509 (open), Copilot SDK mode (engine.copilot-sdk: true) runs each model through a provider for its own wire API, so a Claude sub-agent under a GPT main works there. Please mention SDK mode as the alternative once that lands, and keep the CLI-mode advice.
  3. Add the silent fallback symptom. On AWF v0.28.50 a cross-family sub-agent fails with AWF's Routing model "<model>" to /chat/completions is incompatible … 400, and the main agent then redoes the work on its own model, so the run can succeed with no error in the final output. The checklist should say to look for subagent.failed events in usage/aw_session.jsonl and the audit's "Sub-agent Failed" finding and deviated requests. Please add AWF's Routing model … is incompatible message to the signature list.

Please keep this PR to the #67486 scope.

Copilot AI and others added 2 commits October 10, 2026 21:14
Co-authored-by: SivaKesava1 <11771739+SivaKesava1@users.noreply.github.com>
Co-authored-by: SivaKesava1 <11771739+SivaKesava1@users.noreply.github.com>

Copilot AI commented Oct 10, 2026

Copy link
Copy Markdown
Contributor Author

@copilot Content review of the checklist (in addition to gh-aw-bot's point about running the guidance tests in CI):

  1. The wire API order in step 2 describes main before Select Copilot wire APIs from AWF metadata and surface model mismatches #67487. Select Copilot wire APIs from AWF metadata and surface model mismatches #67487 (open) m...

Updated both checklists for AWF /reflect endpoint metadata and startup mismatch failures, documented the CLI-wide constraint and SDK-mode alternative, and added silent v0.28.50 sub-agent failure signatures/artifacts with contract coverage. Committed as a47b968; make test-guidance and final validation pass.

@SivaKesava1

Copy link
Copy Markdown
Collaborator

@copilot /pr-finisher

Checked a47b968: the checklist now covers AWF /reflect precedence and the startup mismatch, SDK mode as the alternative, and the silent sub-agent fallback. make test-guidance passes (4 tests), and the diff stays within the #67486 scope. Thanks.

One fix before finishing, for the open review thread on test_guidance.py: test-guidance is now part of make test-all, but no workflow in .github/workflows/ runs test-all or test-guidance, so the guard still doesn't run in CI. Please add a step that runs make test-guidance to an existing Ubuntu job in .github/workflows/ci.yml (python3 is preinstalled on GitHub-hosted Ubuntu runners). Run it on changes to .github/skills/agentic-workflows/**, .github/aw/debug-agentic-workflow.md and debug.md if the changes filter supports that, otherwise always. Then reply on that thread. Keep everything else unchanged and within #67486's scope.

Copilot AI and others added 2 commits October 10, 2026 21:38
Co-authored-by: SivaKesava1 <11771739+SivaKesava1@users.noreply.github.com>
Co-authored-by: SivaKesava1 <11771739+SivaKesava1@users.noreply.github.com>
@SivaKesava1

Copy link
Copy Markdown
Collaborator

Checked 14e8202. The guidance tests now run in CI (make test-guidance in ci.yml, filtered on .github/skills/agentic-workflows/**, .github/aw/debug-agentic-workflow.md and debug.md), so we resolved the review thread.

Blocking: the routing line in SKILL.md was removed. This commit deleted the entry that a47b968 added under "follow it directly":

AWF model/endpoint 400s, silent cross-family sub-agent failures, model: auto failures, install-step 404s after a version pin, or questions about version fields: go directly to Model and engine misconfiguration before suggesting model changes.

SKILL.md is now identical to main, so the skill no longer sends these cases to the new checklist. That routing is the core of #67486: the BASF case went wrong because the agent suggested switching to an older model instead of checking the gh-aw version and wire API first. Please restore the line, or explain in a reply why it had to go (for example a size or lint limit) and where these cases are now routed. Please also add a guidance test that asserts SKILL.md routes these signatures to the misconfiguration section, so this can't silently regress.

… in guidance tests

Co-authored-by: SivaKesava1 <11771739+SivaKesava1@users.noreply.github.com>

Copilot AI commented Oct 10, 2026

Copy link
Copy Markdown
Contributor Author

Checked 14e8202. The guidance tests now run in CI (make test-guidance in ci.yml, filtered on .github/skills/agentic-workflows/**, .github/aw/debug-agentic-workflow.md and debug.md), so we ...

The routing line is restored in 0e2132a.

Why it went missing: 14e8202 didn't remove it on purpose. make recompile runs gh aw init, which rebuilds SKILL.md from the embedded template pkg/cli/data/agentic_workflows_skill.md. a47b968 had only edited SKILL.md, so the rebuild dropped the line.

What changed:

  • I added the line to the template and regenerated SKILL.md, so it now stays after init or make recompile.
  • test_routing_and_evidence_links now checks both SKILL.md and the template. Each must have a line linking to ../../aw/debug-agentic-workflow.md#model-and-engine-misconfiguration that lists AWF model/endpoint 400s, silent cross-family sub-agent failures, `model: auto` failures, install-step 404s after a version pin, and version fields.

make test-guidance passes with the fix and fails without it. No other files changed.

@pelikhan
pelikhan merged commit 9ded002 into main Oct 10, 2026
33 of 34 checks passed
@pelikhan
pelikhan deleted the copilot/add-checklist-for-misconfiguration branch October 10, 2026 23:48
github-actions Bot added a commit that referenced this pull request Oct 11, 2026
memory.md had grown to 407 lines, over the 400-line instruction-file
cap, after recent repo-memory JSON-schema-validation changes (#67439).
Extract the self-contained "standalone ledger (experimental)" section
into a new ledger.md, replacing it with a short pointer. Update the
cross-reference in memory-stateful-patterns.md to link directly to
ledger.md instead of an anchor inside memory.md.

debug-agentic-workflow.md is also over the cap (432 lines) after
recent model/engine-misconfiguration diagnostics (#67488), but its
checklist content is pinned inline by a regression test
(.github/skills/agentic-workflows/tests/test_guidance.py) that also
requires it mirrored in root debug.md, so it is left as a documented
exception rather than split.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

agentic-workflows skill: add a checklist for model and engine misconfiguration

5 participants