Skip to content

Cut AI Moderator token usage by pre-fetching capped moderation context - #52891

Merged
pelikhan merged 2 commits into
mainfrom
copilot/deep-report-investigate-token-usage
Aug 15, 2026
Merged

pelikhan merged 2 commits into
mainfrom
copilot/deep-report-investigate-token-usage

Conversation

Copilot AI commented Aug 15, 2026 •

Copy link
Copy Markdown
Contributor

The AI Moderator workflow was flagged for abnormally high token usage. A 30-run / 14-day sample confirms the lead: 3.0M tokens across 30 runs (~100k/run) versus a ~26k fleet average, with 47–109 GitHub API calls per run.

Root cause

Audit of run 31867160774: 11 model requests × ~12–15k input tokens each. The workflow runs in gh-proxy mode (no MCP tool schemas), so cost is turns × ambient context. The prompt explicitly told the agent to fetch all issue/comment/PR content itself — including an uncapped pull_request_read get_diff — so every fetch/analysis step added another full-context turn.

Changes to .github/workflows/ai-moderator.md

  • pre-agent-steps DataOps prefetch — deterministic gh api calls write a compact /tmp/gh-aw/agent/moderation-context.json (event, actor, item, comment; bodies clipped to 6000 chars) and /tmp/gh-aw/agent/pr-diff.patch capped at 200 lines. Same pattern as shared/pr-diff-data-fetch.md used by PR Code Quality Reviewer.
  • Prompt rewrite — Context section points at the pre-fetched files and forbids re-fetching; the PR branch of Actions reads the capped patch instead of get_diff; the spam log is read in the same turn as the context files.
  • Turn budget — explicit instruction to keep the run to one read, one analysis, one safe output.
## Context

The content to moderate has already been fetched for you — **do not call GitHub APIs,
`gh`, `issue_read`, or `pull_request_read` to fetch it again**:

- `/tmp/gh-aw/agent/moderation-context.json` — ... Bodies are the original unsanitized
  user input, truncated to 6000 characters.
- `/tmp/gh-aw/agent/pr-diff.patch` — the pull request diff, capped at the first 200 lines.

Bodies remain unsanitized original content, preserving detection fidelity; only length is bounded.

Notes for reviewers

  • ai-moderator.md carries redirect: githubnext/agentics/workflows/ai-moderator.md@main, so a future gh aw update would overwrite this. The same change should be landed upstream to make it stick.
  • No measurement of the post-change cost is possible from here; worth re-sampling after a few runs, or gating behind an experiments: variant if the reviewer prefers an A/B before promoting.
  • The model tier was left alone — gh aw audit suggests a downgrade is plausible, but that trades detection quality and should be measured separately.

Co-authored-by: pelikhan <4175913+pelikhan@users.noreply.github.com>
@github-actions

Copy link
Copy Markdown
Contributor

Thanks for diving into the AI Moderator token usage investigation 🔬! This draft PR is well-scoped and properly linked to #52739. The task checklist is clear:

  • ✅ Sampling and auditing phases complete
  • 📝 Remaining work: deterministic pre-fetch, prompt slimming, and lock-file recompilation

Looks like the investigation phase is wrapping up nicely. Good luck with the optimization pass!

Generated by ✅ Contribution Check · auto · 53.5 AIC · ⌖ 4.72 AIC · ⊞ 9.1K · ◷

Copilot AI changed the title [WIP] Investigate high token usage in AI Moderator workflow Cut AI Moderator token usage by pre-fetching capped moderation context Aug 15, 2026
Copilot AI requested a review from pelikhan August 15, 2026 13:12
@pelikhan
pelikhan marked this pull request as ready for review August 15, 2026 13:20
Copilot AI balanced review requested due to automatic review settings August 15, 2026 13:20
@pelikhan
pelikhan merged commit 4d3e661 into main Aug 15, 2026
@pelikhan
pelikhan deleted the copilot/deep-report-investigate-token-usage branch August 15, 2026 13:20

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Reduces AI Moderator token usage by moving GitHub context retrieval into a bounded deterministic prefetch step.

Changes:

  • Prefetches capped issue, comment, and PR diff context.
  • Rewrites moderation instructions to use prefetched files with fewer turns.
  • Regenerates the compiled workflow.
Show a summary per file
File Description
.github/workflows/ai-moderator.md Adds context prefetching and revises the agent prompt.
.github/workflows/ai-moderator.lock.yml Incorporates the generated prefetch step.

Review details

💡 Add a code-review agent skill for context-aware, tailored reviews. Learn more in the docs.

Suppressed comments (2)

.github/workflows/ai-moderator.md:92

  • Swallowing a comment-fetch failure leaves comment null while the issue_comment run continues, so the agent may assess the parent issue instead of the triggering comment and emit an incorrect action. Fail the prefetch when the required comment cannot be retrieved.
        gh api "repos/$EXPR_GITHUB_REPOSITORY/issues/comments/$COMMENT_ID" > "$RAW_COMMENT" || echo '{}' > "$RAW_COMMENT"

.github/workflows/ai-moderator.md:96

  • || true also hides genuine authentication, network, and API failures, producing an empty patch that the agent can treat as a clean diff and label ai-inspected. Capture the command output before truncating so a real fetch failure stops the run without the head pipeline's SIGPIPE concern.
      if [ -n "${PR_NUMBER:-}" ]; then
        { gh pr diff "$PR_NUMBER" --repo "$EXPR_GITHUB_REPOSITORY" || true; } \
          | head -n "$DIFF_MAX_LINES" > /tmp/gh-aw/agent/pr-diff.patch
  • Files reviewed: 2/2 changed files
  • Comments generated: 3
  • Review effort level: Balanced

agent:
runtime: gvisor
sudo: false
pre-agent-steps:
echo '{}' > "$RAW_COMMENT"
ITEM_NUMBER="${ISSUE_NUMBER:-${PR_NUMBER:-}}"
if [ -n "$ITEM_NUMBER" ]; then
gh api "repos/$EXPR_GITHUB_REPOSITORY/issues/$ITEM_NUMBER" > "$RAW_ISSUE" || echo '{}' > "$RAW_ISSUE"
PR_NUMBER: ${{ github.event.pull_request.number }}
COMMENT_ID: ${{ github.event.comment.id }}
BODY_MAX_CHARS: "6000"
DIFF_MAX_LINES: "200"
@github-actions

Copy link
Copy Markdown
Contributor

🎉 This pull request is included in a new release.

Release: v0.86.3

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[deep-report] Investigate abnormally high token usage in AI Moderator workflow (~4.7x fleet average)

3 participants