Executive Summary
- Runs inspected: 2 audited (Matt Pocock Skills Reviewer, Impeccable Skills Reviewer); 60 runs listed in the 24h window.
- **Reduced (redacted) the per-run artifact directories (
/tmp/gh-aw/aw-mcp/logs/run-*) are permission-denied in this sandbox, so first-request text could not be extracted. The Python analysis was not run and no char/line metrics exist. Figures below come from audit (prompt_analysis, ambient_context, working_set).
- Ambient first-request context is ~5.4k–7.4k input tokens. WSRF is ~1.0–1.1, so static context is not being rebuilt per turn. The main cost is cache-read volume across many turns (87 and 53 requests per run), not ambient size.
- The largest compiled prompts are 19.9k chars (Matt Pocock) and 14.4k chars (Impeccable).
Highest-Leverage Changes
impeccable-skills-reviewer.md: move the process and review-mode sections (lines 91–153) into an on-demand skill. Impact: medium. Needs manual review.
deep-report.md (18.8 KB): the large inline report template (lines ~318–367) belongs in a ## skill: block or shared import. Impact: medium. Needs manual review.
linter-miner.md (11.6 KB): the inline File 1/2/3 templates and 3 inline agents (lines 178–272) can be loaded on demand. Impact: medium. Needs manual review.
- Skills reviewers: cap tool turns and requests (87 requests, 226 AIC). Cost is driven by turn count rather than prompt size.
mattpocock-skills-reviewer.md is retry-blocked, so apply this to the Impeccable reviewer only. Impact: medium.
Blocked (closed PR #65568 in the last 14 days): mattpocock-skills-reviewer.md, pr-code-quality-reviewer.md, test-quality-sentinel.md, daily-ambient-context-optimizer.md. These were excluded.
CI-Validation Checklist for Implementing Agents
Key Metrics
| Metric |
Value |
| Sampled runs |
2 (audit-only) |
| Distinct workflows |
2 |
| Median chars |
n/a (artifacts unreadable) |
| P95 chars |
n/a |
| Largest sampled request |
prompt.txt 19,937 chars (Matt Pocock) |
| Merged optimizer PRs (7d) |
0 |
| Closed optimizer PRs (7d) |
1 |
| Optimizer PR close-rate (7d) |
n/a (<3 settled PRs; auto_pause=false) |
| WSRF (audited runs) |
1.10 (peak 7,378 tok); 1.02 (peak 5,420 tok) |
Per-Run First-Request Metrics
| Run |
Workflow |
AIC |
Requests |
Ambient input tok |
prompt.txt chars |
WSRF |
| 37215431563 |
Matt Pocock Skills Reviewer |
226.4 |
87 |
7,378 |
19,937 |
1.10 |
| 37215431740 |
Impeccable Skills Reviewer |
115.6 |
53 |
5,420 |
14,377 |
1.02 |
Repeated Ambient Context Signals
- Both reviewers import shared PR-review, reporting, otlp and pr-diff-data-fetch modules.
deep-report and linter-miner also pull in shared/reporting.md and shared/otlp.md. Together these are the common repeated fragments.
deep-report, linter-miner and Impeccable already set tools.cli-proxy: true. deep-report sets explicit GitHub toolsets (default, actions, discussions, search). Impeccable's toolsets were not confirmed.
- Raw
gh aw instruction wording was not checked, because the first-request text was unavailable.
Deterministic Analysis Output
Not produced: no first-request source was readable. The metrics above are the audit's compact values only.
Recommendations by Category
Workflow Markdown
- Move the large output templates in
deep-report.md and the file templates in linter-miner.md out of the main prompt body. Impact: medium. Needs manual review.
- In
impeccable-skills-reviewer.md, turn the review-mode list into on-demand context. Impact: medium.
- Add explicit tool-turn budgets to the skills reviewers. Impact: medium.
Skills
- Impeccable already imports a pinned upstream skill. Load only the mode that is needed, not the whole skill. Impact: low–medium.
Agents
linter-miner defines 3 inline agents (discussion-miner, code-pattern-scanner, linter-writer). Check whether each is invoked on every run, and drop or lazy-load any that rarely are. Impact: low–medium. Needs manual review.
References
Generated by 🌫️ Daily Ambient Context Optimizer · copilot · auto · 24.5 AIC · ⌖ 15.8 AIC · ⊞ 12K · ◷
Executive Summary
/tmp/gh-aw/aw-mcp/logs/run-*) are permission-denied in this sandbox, so first-request text could not be extracted. The Python analysis was not run and no char/line metrics exist. Figures below come fromaudit(prompt_analysis,ambient_context,working_set).Highest-Leverage Changes
impeccable-skills-reviewer.md: move the process and review-mode sections (lines 91–153) into an on-demand skill. Impact: medium. Needs manual review.deep-report.md(18.8 KB): the large inline report template (lines ~318–367) belongs in a## skill:block or shared import. Impact: medium. Needs manual review.linter-miner.md(11.6 KB): the inline File 1/2/3 templates and 3 inline agents (lines 178–272) can be loaded on demand. Impact: medium. Needs manual review.mattpocock-skills-reviewer.mdis retry-blocked, so apply this to the Impeccable reviewer only. Impact: medium.Blocked (closed PR
#65568in the last 14 days):mattpocock-skills-reviewer.md,pr-code-quality-reviewer.md,test-quality-sentinel.md,daily-ambient-context-optimizer.md. These were excluded.CI-Validation Checklist for Implementing Agents
make recompilefor every modified.github/workflows/*.mdfile; zero errors requiredmake agent-report-progressbefore the final commit and confirm it passesblocked_filesin/tmp/gh-aw/ambient-context/closed-pr-targets.json; do not re-attempt files from closed ambient-context PRs in the last 14 days.lock.ymlchanges in the PR bodyKey Metrics
Per-Run First-Request Metrics
Repeated Ambient Context Signals
deep-reportandlinter-mineralso pull inshared/reporting.mdandshared/otlp.md. Together these are the common repeated fragments.deep-report,linter-minerand Impeccable already settools.cli-proxy: true.deep-reportsets explicit GitHub toolsets (default, actions, discussions, search). Impeccable's toolsets were not confirmed.gh awinstruction wording was not checked, because the first-request text was unavailable.Deterministic Analysis Output
Not produced: no first-request source was readable. The metrics above are the audit's compact values only.
Recommendations by Category
Workflow Markdown
deep-report.mdand the file templates inlinter-miner.mdout of the main prompt body. Impact: medium. Needs manual review.impeccable-skills-reviewer.md, turn the review-mode list into on-demand context. Impact: medium.Skills
Agents
linter-minerdefines 3 inline agents (discussion-miner,code-pattern-scanner,linter-writer). Check whether each is invoked on every run, and drop or lazy-load any that rarely are. Impact: low–medium. Needs manual review.References