You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Test Quality Sentinel (.github/workflows/test-quality-sentinel.md). It is the highest-AIC workflow not excluded: Matt Pocock Skills Reviewer and Impeccable Skills Reviewer were optimized within the last 14 days, so they are skipped.
Analysis period
7 days to 2026-10-05, 5 runs, all audited from all-runs.json. 4 runs succeeded and 1 failed.
Cost profile
Metric
Value
Total AIC
116.32
Avg AIC/run
23.26
Median AIC/run
6.91
Raw tokens (total)
84,833
Avg turns/run
4.2 (21 total)
Action minutes
54
The 4 healthy runs cost 4.1–13.5 AIC each, 29.5 AIC combined. One failed run cost 86.8 AIC (75% of the total).
Ranked recommendations
1. Stop the runaway continuation loop (est. savings ~15–20 AIC/run averaged, ~80 AIC per incident)
Evidence:§37322950862 used 16 turns, 83 model invocations and 18.3 min. Its AIC was 86.8, versus a ~7 AIC median.
The run had already posted add_comment and submit_pull_request_review at 14:28 but still ended as failure (agent_logic).
Working-set rebuild factor was 1.50, with 3,153 excess tokens. Healthy runs sit at ~1.00.
Action:
Lower engine.max-continuations from 15 to 3–4. Healthy runs need 1–2 turns.
Add a prompt line to the Step 8 closing text: "After the single add_comment and single review/noop call, stop immediately. Do not re-read data or re-submit."
Note that add-comment and submit-pull-request-review are capped at max 1 each. Retries beyond that are pure waste.
2. Drop the echo:* and git diff:* bash allowlist entries if unused (est. 0–1 AIC/run, low risk)
All data is pre-fetched, and the prompt says to use cat/grep on /tmp/gh-aw/agent/*.
The github: mode: gh-proxy and cli-proxy tools show only 2–8 GitHub API calls per run.
Action: tool-removal rule: audit only. Keep git diff since it is used in 1 run (8 API calls); consider echo removal only after checking agent logs.
3. Read pre-fetched files in one call (est. 1–2 AIC/run)
Steps 1 and 3 list about 8 separate files (pr-meta.json, test-files.txt, test-diff.txt, go-test-stats.txt and others).
Action: instruct "run one cat over all listed files in a single bash call" instead of per-file reads.
4. Trim prompt (est. 1–2 AIC/run)
The prompt body is ~12 KB of the 26 KB source.
The rubric and red-flag lists are stated in Step 2, Step 4 and the act-vs-noop table.
The infrastructure-only conditions are repeated in Steps 2 and 6.
Action: keep one canonical copy and reference it.
Structural optimizations
No inline sub-agent recommended. The review requires synthesis across tests, and the healthy runs are already cheap.
No shared setup extraction. Only Step 1 does setup.
Caveats
The cause of the failed run's loop is inferred from run metrics (turns, invocations, rebuild factor, agent_logic), not from a full agent log read.
The sample is 5 runs, and just 1 failed. The savings estimate for rejig docs #1 depends on the failure rate (~20% here).
The model-size experiment (haiku-4.5 vs sonnet-5) varies cost per run, which can explain part of the spread among healthy runs.
The workflow has no source: entry, so it can be edited directly. Recompile after any change.
Target workflow
Test Quality Sentinel (
.github/workflows/test-quality-sentinel.md). It is the highest-AIC workflow not excluded: Matt Pocock Skills Reviewer and Impeccable Skills Reviewer were optimized within the last 14 days, so they are skipped.Analysis period
7 days to 2026-10-05, 5 runs, all audited from
all-runs.json. 4 runs succeeded and 1 failed.Cost profile
The 4 healthy runs cost 4.1–13.5 AIC each, 29.5 AIC combined. One failed run cost 86.8 AIC (75% of the total).
Ranked recommendations
1. Stop the runaway continuation loop (est. savings ~15–20 AIC/run averaged, ~80 AIC per incident)
add_commentandsubmit_pull_request_reviewat 14:28 but still ended asfailure(agent_logic).engine.max-continuationsfrom 15 to 3–4. Healthy runs need 1–2 turns.add_commentand single review/noop call, stop immediately. Do not re-read data or re-submit."add-commentandsubmit-pull-request-revieware capped at max 1 each. Retries beyond that are pure waste.2. Drop the
echo:*andgit diff:*bash allowlist entries if unused (est. 0–1 AIC/run, low risk)cat/grepon/tmp/gh-aw/agent/*.github: mode: gh-proxyandcli-proxytools show only 2–8 GitHub API calls per run.git diffsince it is used in 1 run (8 API calls); considerechoremoval only after checking agent logs.3. Read pre-fetched files in one call (est. 1–2 AIC/run)
pr-meta.json,test-files.txt,test-diff.txt,go-test-stats.txtand others).catover all listed files in a single bash call" instead of per-file reads.4. Trim prompt (est. 1–2 AIC/run)
Structural optimizations
Caveats
agent_logic), not from a full agent log read.haiku-4.5vssonnet-5) varies cost per run, which can explain part of the spread among healthy runs.source:entry, so it can be edited directly. Recompile after any change.References: §37322950862, §37327387777, §37325263100