Skip to content

Fix prompt-exhaustion misclassification and restore Avenger's bash tool access - #67411

Merged
pelikhan merged 4 commits into
mainfrom
pelikhan-issue-66817-assess-scheduled-prompt-exhaustion-1aeb98
Oct 10, 2026
Merged

pelikhan merged 4 commits into
mainfrom
pelikhan-issue-66817-assess-scheduled-prompt-exhaustion-1aeb98

Conversation

@pelikhan

@pelikhan pelikhan commented Oct 10, 2026 •

Copy link
Copy Markdown
Collaborator

Root cause #1 — misclassification (fixed)

buildEmptyOutputOutcome() in actions/setup/js/empty_output_outcome.cjs computed the failureCause for a missing_terminal_safe_output / invalid_safe_outputs incompletion with:

const failureCause = isPromptExhaustion ? "prompt_exhaustion" : isRequestRejection ? "request_rejection" : isEngineOutage ? "engine_outage" : "prompt_exhaustion";

The final fallback defaulted to "prompt_exhaustion" whenever none of the three evidence-backed buckets matched — i.e. no effective_tokens_limit_exceeded/invocation_cap_exceeded category, no 4xx/request-rejection signal, and no 5xx/engine-driver-failure signal. Absence of evidence was being treated as evidence of prompt exhaustion.

This is confirmed across the duplicate source issues clustered under #66817 (e.g. #67330, #67358): all show Failure classification: prompt_exhaustion with Retry attempts observed: 0 and no token/cache-miss/invocation-cap diagnostics — exactly the "no evidence" fallback path, not real prompt exhaustion.

Fix: Added a new unknown failure cause (EMPTY_OUTPUT_FAILURE_CAUSES.unknown = "finished without a clear failure cause") and changed the fallback to use it instead of guessing prompt_exhaustion. This flows into buildFailureIssueTitle() (handle_agent_failure.cjs) and the failure_cause: issue marker used for dedup/search.

Root cause #2 — Avenger's specific silent "no terminal output" failures (fixed)

Review correctly flagged that fixing the classification label alone did not address Avenger's explicit acceptance criterion: Avenger was one of the misclassified workflows, and its real runs were failing for a concrete, separate reason, not budget/token exhaustion.

Downloaded and inspected real Avenger run logs (e.g. run 37994092775, issue #67282). Evidence: a single turn, ~16s codex process duration, usage {input_tokens:78281, output_tokens:245} (nowhere near exhaustion), and the agent's own message:

"Unable to proceed: this session exposes no shell tool, CI job/log reader, or safe-output tools... No changes or PR were created."

Traced to pkg/workflow/codex_config.go (buildNativeConfig): the Codex engine unconditionally sets features.shell_tool = false (and code_mode/code_mode_only = false) whenever engine.model-provider resolves to github — regardless of whether tools.bash is enabled. This is an intentional, documented compatibility workaround (PRs #65654, #66179) because "the Copilot compatibility adapter does not support the Codex exec custom tool" — but Codex has no MCP-routed fallback for bash (codexNativeServerDefaults never maps bash/cli-proxy to an MCP server), so for any workflow combining engine: codex + model-provider: github + tools.bash, bash is declared but completely non-functional at runtime, silently stripping tool access. Avenger declared tools.bash: ["*"] while using exactly this combination.

Fix: switched Avenger (.github/workflows/avenger.md) from engine.id: codex to engine.id: copilot, which enforces tools.bash natively (--allow-tool shell) with no GitHub-provider restriction — no prompt/model/permission changes. Confirmed via recompile that avenger.lock.yml no longer disables shell_tool and now runs through copilot_harness.cjs with --allow-tool shell.

Added pkg/workflow/avenger_engine_regression_test.go — a regression test that reads Avenger's live frontmatter and fails if it ever reverts to an engine that can't enforce a bash allowlist while model-provider: github + tools.bash remain enabled. Verified the test fails when engine.id is reverted to codex and passes with copilot.

Scope note: this is Avenger-specific. A broader survey found ~20 other workflows combine engine: codex + GitHub-provider inference + tools.bash, which share the same latent gap — out of scope for this PR but worth a follow-up.

Files changed

  • actions/setup/js/empty_output_outcome.cjs — classification fix
  • actions/setup/js/empty_output_outcome.test.cjs, handle_agent_failure.test.cjs, collect_ndjson_output.test.cjs — updated expectations/fixtures
  • .github/workflows/avenger.md, .github/workflows/avenger.lock.yml — engine switch (codex → copilot) to restore bash/tool access
  • pkg/workflow/avenger_engine_regression_test.go — new regression guard

Validation

  • make agent-report-progress (full gate): Go lint ✓, JS lint ✓, schema freshness ✓, impacted tests ✓, 614/614 JS tests passed
  • make recompile: 334/334 workflows compiled successfully; confirmed avenger.lock.yml no longer contains "shell_tool":false and now invokes copilot_harness.cjs --allow-tool shell
  • New regression test (TestAvengerUsesBashCapableEngine) added and manually verified to fail when the engine is reverted to codex
  • Limitation: this fix is validated against historical run-log evidence, static config, and tests — not a live re-run, since triggering scheduled workflows is out of scope for this change. The next real Avenger CI-fix-needed run will be the first live confirmation.

Fixes #66817

buildEmptyOutputOutcome() fell back to "prompt_exhaustion" whenever no
PROMPT_EXHAUSTION_ERROR_CATEGORIES, REQUEST_REJECTION_ERROR_CATEGORIES,
or 5xx engine-outage evidence was found. Absence of evidence is not
evidence of token/invocation-limit exhaustion, so most
missing_terminal_safe_output incompletions (no matched category, 0
retries observed) were mislabeled as prompt_exhaustion. This inflated
the prompt_exhaustion failure-issue cluster (#66817) across
unrelated workflows such as Avenger.

Classify the unmatched case as a new "unknown" failure cause instead of
defaulting to prompt_exhaustion, with a matching issue-title mapping.

Fixes #66817

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@pelikhan
pelikhan marked this pull request as ready for review October 10, 2026 12:05
Copilot AI balanced review requested due to automatic review settings October 10, 2026 12:05
…assess-scheduled-prompt-exhaustion-1aeb98

# Conflicts:
#	actions/setup/js/empty_output_outcome.cjs
#	actions/setup/js/handle_agent_failure.test.cjs

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

Declaring Fixes #66817 would close an issue whose failure-prevention scope and acceptance criteria remain unaddressed.

1 open finding
What changed in this PR

Reclassifies terminal-output failures lacking evidence from prompt_exhaustion to unknown.

Changes:

  • Adds the unknown failure cause and issue title.
  • Updates collector and failure-handler test expectations.
  • Preserves evidence-backed classifications.
File Description
actions/​setup/​js/​empty_output_outcome.cjs Adds the unknown fallback classification.
actions/​setup/​js/​empty_output_outcome.test.cjs Tests unknown classification.
actions/​setup/​js/​handle_agent_failure.test.cjs Tests the unknown-cause issue title.
actions/​setup/​js/​collect_ndjson_output.test.cjs Updates collector fixture expectations.

🧠 Review effort: Balanced


💡 Add a code-review agent skill for context-aware, tailored reviews. Learn more in the docs.

Comment thread actions/setup/js/empty_output_outcome.cjs
@pelikhan pelikhan changed the title Fix default failure-cause misclassification as prompt_exhaustion Fix prompt-exhaustion misclassification and restore Avenger's bash tool access Oct 10, 2026
pelikhan and others added 2 commits October 10, 2026 05:29
Root cause: Codex engine's buildNativeConfig() unconditionally sets
features.shell_tool=false (plus code_mode/code_mode_only=false) whenever
engine.model-provider resolves to "github", regardless of whether
tools.bash is enabled. Codex has no MCP-routed fallback for bash, so
Avenger's tools.bash: ["*"] was silently inert at runtime despite being
declared. Real run-log evidence (run 37994092775) shows the agent itself
reporting zero available tools and exiting cleanly after a single short
turn — not token/prompt exhaustion, which is what the "prompt_exhaustion"
classification had been masking.

Fix: switch avenger.md's engine.id from codex to copilot, which enforces
tools.bash natively via --allow-tool shell and has no GitHub-provider
shell-disabling restriction. Model, permissions, sandbox, and prompt are
unchanged. Recompiled avenger.lock.yml: confirmed it no longer disables
shell_tool and now runs through copilot_harness.cjs with --allow-tool
shell.

Added pkg/workflow/avenger_engine_regression_test.go to guard against
regressing Avenger back onto an engine that cannot enforce a bash
allowlist while model-provider: github and tools.bash remain enabled.
Verified the test fails when engine.id is reverted to codex.

Addresses the Avenger-specific acceptance criterion in #66817 that the
classification-only fix did not cover.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
CI's lint-js job runs 'npx prettier --check' against actions/setup/js,
which was failing on a single formatting nit in a for-loop header
(trailing space before the closing paren) unrelated to this PR's
other changes. Applied 'prettier --write' to restore compliant
formatting; verified the exact CI command now passes.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[AW Top 10] 02 Reduce prompt exhaustion across scheduled agent workflows

2 participants