Summary
go-fan (Go Fan — Daily Go Module Reviewer) is the only workflow in the current evidence window that succeeded yet logged an avg_aic an order of magnitude above every other successful run. This proposes a generation control/prompt_tightening experiment that bounds its open-ended Step 3 research scope, aiming to cut agent turns (execution-step-count) without degrading review completeness.
Observation
In per-workflow-summary.json, Go Fan's one recorded run succeeded (failure_count=0) but logged avg_aic=186.528795 — 6–17x higher than every other successful run in the same slice (daily-experiment-report=21.27, Daily Syntax Error Quality=13.06, Issue Monster=11.19, aw-failure-investigator=0). Reading .github/workflows/go-fan.md, Step 3 ("Research the Module") instructs the agent to open-endedly "explore the module's repository" — README, releases/changelog, and "popular usage examples in issues/discussions" — with no bound on how many tool calls/turns that exploration may take.
Hypothesis
Mechanism: Bounding Step 3 to a fixed, prioritized set of research actions (README quickstart section, latest 3 releases/changelog entries, up to 5 targeted usage-pattern search results) removes the unconstrained issue/discussion browsing that currently has no stopping condition, cutting the number of LLM turns spent on research before the agent moves to Step 4 (Serena analysis) and beyond.
Expected direction: decrease in execution-step-count (total LLM request count).
Minimum effect: >=3 absolute count decrease in execution-step-count, averaged per run.
Candidate Mutation
- primary_dimension:
generation control
- subtype:
prompt_tightening
- Control: Step 3 keeps its existing open-ended instructions — explore the repository's README, releases/changelog, and issue/discussion usage examples with no bound — unchanged from today.
- Candidate: Step 3's "3.1 GitHub Repository" and "3.3 Recent Updates" subsections are replaced with a bounded version: fetch only the README's primary usage section, the latest 3 releases/changelog entries, and up to 5 targeted usage-pattern search results, explicitly skipping further issue/discussion browsing. "3.2 Documentation" is untouched in both branches.
Experiment
graders:
execution-step-count: {}
tool-failure-count: {}
evals:
- id: module_review_completeness
question: "Did the agent produce a complete module review — covering the module's overview, current project usage, research findings, and at least one improvement opportunity — saved to scratchpad/mods/<module>.md, plus a matching issue?"
experiments:
prompt_tightening_v1:
variants: [control, candidate]
description: "Bound Step 3 (Research the Module) to a fixed, prioritized set of research actions (README quickstart, latest 3 releases/changelog entries, up to 5 targeted usage-pattern search results) instead of open-ended repository/issue/discussion exploration, to cut the number of agent turns spent on unconstrained research."
hypothesis: "H0: No meaningful difference in execution-step-count between control and candidate. H1: Bounding Step 3's research scope decreases execution-step-count without lowering the module_review_completeness eval pass rate."
metric: "grader:execution-step-count"
guardrail_metrics:
- name: "eval:module_review_completeness"
threshold: ">=0.90"
- name: "grader:tool-failure-count"
threshold: "<=5"
min_samples: 20
analysis_type: mann_whitney
decision:
minimum_effect: 3
regression_tolerance: 3
confidence: 0.95
tags: ["harness_dimension:generation control", "harness_subtype:prompt_tightening"]
Guardrails
| Guardrail |
Threshold |
Why it protects against regression |
eval:module_review_completeness |
>=0.90 |
Since execution-step-count (a cost/efficiency proxy) is the primary metric, this eval verifies the bounded research scope still yields a complete module review (overview, usage, findings, >=1 improvement opportunity, issue + scratchpad summary) — protecting correctness/completeness equivalence against the candidate. |
grader:tool-failure-count |
<=5 |
Guards against the tightened instructions causing new tool-call confusion or failures relative to today's baseline behavior. |
Expected Economics
Go Fan runs on a cron: daily around 7:00 on weekdays schedule (~5 runs/week). At that cadence, reaching min_samples: 20 per variant (40 total observations) takes roughly 8 weeks — within the "few months" bound for sparse-traffic eligibility, and the workflow is not dispatch-only. If the hypothesis holds, the candidate variant is expected to reduce execution-step-count (and by extension avg_aic, currently 186.53 for the one recorded run vs a fleet successful-run median in the 11–21 range) by bounding what is today unconstrained issue/discussion exploration in Step 3, while the module_review_completeness eval guards against a shallower, lower-quality review.
Validation
Both ./gh-aw compile go-fan --strict and ./gh-aw compile go-fan --strict --validate exited 0 (with an expected informational ⚠ Using experimental feature: graders warning, not an error). git diff --stat touched only .github/workflows/go-fan.md and its matching .github/workflows/go-fan.lock.yml. The compiled lock file confirms prompt_tightening_v1 is wired into the activation job's pick-experiment step and into the GH_AW_EXPERIMENTS_PROMPT_TIGHTENING_V1 env var used by the agent job templating.
Interpretation
Applying this patch only starts the experiment; no decision is interpreted here. Once min_samples: 20 is reached, gh aw experiments analyze go-fan computes the deterministic EXTEND/PROMOTE/REJECT/INCONCLUSIVE decision unchanged — this workflow never recomputes, reinterprets, or overrides that decision, and any eventual PROMOTE still requires a separate, human-reviewed change through the existing daily-experiment-report deterministic decision engine. This workflow never merges anything itself.
Rollback
git revert the commit that applies this patch. No other files depend on this change — the experiments:/graders:/evals: blocks and the {{#if}} conditional are fully self-contained within go-fan.md, and reverting regenerates an unchanged go-fan.lock.yml via gh aw compile go-fan.
Manual Patch (apply by hand)
diff --git a/.github/workflows/go-fan.md b/.github/workflows/go-fan.md
index f515da9..a921642 100644
--- a/.github/workflows/go-fan.md
+++ b/.github/workflows/go-fan.md
@@ -31,6 +31,34 @@ engine: claude
name: Go Fan
strict: true
timeout-minutes: 30
+
+graders:
+ execution-step-count: {}
+ tool-failure-count: {}
+
+evals:
+ - id: module_review_completeness
+ question: "Did the agent produce a complete module review — covering the module's overview, current project usage, research findings, and at least one improvement opportunity — saved to scratchpad/mods/<module>.md, plus a matching issue?"
+
+experiments:
+ prompt_tightening_v1:
+ variants: [control, candidate]
+ description: "Bound Step 3 (Research the Module) to a fixed, prioritized set of research actions (README quickstart, latest 3 releases/changelog entries, up to 5 targeted usage-pattern search results) instead of open-ended repository/issue/discussion exploration, to cut the number of agent turns spent on unconstrained research."
+ hypothesis: "H0: No meaningful difference in execution-step-count between control and candidate. H1: Bounding Step 3's research scope decreases execution-step-count without lowering the module_review_completeness eval pass rate."
+ metric: "grader:execution-step-count"
+ guardrail_metrics:
+ - name: "eval:module_review_completeness"
+ threshold: ">=0.90"
+ - name: "grader:tool-failure-count"
+ threshold: "<=5"
+ min_samples: 20
+ analysis_type: mann_whitney
+ decision:
+ minimum_effect: 3
+ regression_tolerance: 3
+ confidence: 0.95
+ tags: ["harness_dimension:generation control", "harness_subtype:prompt_tightening"]
+
tools:
bash:
- cat go.mod
@@ -131,6 +159,29 @@ From the sorted list (most recent first):
## Step 3: Research the Module
+{{#if experiments.prompt_tightening_v1 == 'candidate' }}
+For the selected module, research it with a **bounded, prioritized** scope — do not open-endedly browse the repository:
+
+### 3.1 GitHub Repository (bounded)
+Use GitHub tools to fetch only:
+- The README's primary usage/quickstart section
+- The latest 3 releases (or changelog entries if no releases)
+- Up to 5 search results for usage examples in issues/discussions — skip further issue/discussion browsing beyond this
+
+### 3.2 Documentation
+Note key features and API patterns:
+- Core APIs and their purposes
+- Common usage patterns
+- Performance considerations
+- Recommended configurations
+
+### 3.3 Recent Updates
+From the releases/changelog already fetched in 3.1, note:
+- New features in recent releases
+- Breaking changes
+- Deprecations
+- Security advisories
+{{#else}}
For the selected module, research its:
### 3.1 GitHub Repository
@@ -153,6 +204,7 @@ Check for:
- Breaking changes
- Deprecations
- Security advisories
+{{#endif}}
## Step 4: Analyze Project Usage with Serena
Application Plan
- Save the diff above as
proposal.patch.
git apply proposal.patch
gh aw compile go-fan --strict --validate
- Open a PR manually if compilation succeeds.
Generated by 🧫 Daily Harness Experiment Proposer · claude · agent · 203.7 AIC · ⊞ 7.3K · ◷
Summary
go-fan(Go Fan — Daily Go Module Reviewer) is the only workflow in the current evidence window that succeeded yet logged anavg_aican order of magnitude above every other successful run. This proposes ageneration control/prompt_tighteningexperiment that bounds its open-ended Step 3 research scope, aiming to cut agent turns (execution-step-count) without degrading review completeness.Observation
In
per-workflow-summary.json, Go Fan's one recorded run succeeded (failure_count=0) but loggedavg_aic=186.528795— 6–17x higher than every other successful run in the same slice (daily-experiment-report=21.27,Daily Syntax Error Quality=13.06,Issue Monster=11.19,aw-failure-investigator=0). Reading.github/workflows/go-fan.md, Step 3 ("Research the Module") instructs the agent to open-endedly "explore the module's repository" — README, releases/changelog, and "popular usage examples in issues/discussions" — with no bound on how many tool calls/turns that exploration may take.Hypothesis
Mechanism: Bounding Step 3 to a fixed, prioritized set of research actions (README quickstart section, latest 3 releases/changelog entries, up to 5 targeted usage-pattern search results) removes the unconstrained issue/discussion browsing that currently has no stopping condition, cutting the number of LLM turns spent on research before the agent moves to Step 4 (Serena analysis) and beyond.
Expected direction: decrease in
execution-step-count(total LLM request count).Minimum effect: >=3 absolute count decrease in
execution-step-count, averaged per run.Candidate Mutation
generation controlprompt_tighteningExperiment
Guardrails
eval:module_review_completeness>=0.90execution-step-count(a cost/efficiency proxy) is the primary metric, this eval verifies the bounded research scope still yields a complete module review (overview, usage, findings, >=1 improvement opportunity, issue + scratchpad summary) — protecting correctness/completeness equivalence against the candidate.grader:tool-failure-count<=5Expected Economics
Go Fan runs on a
cron: daily around 7:00 on weekdaysschedule (~5 runs/week). At that cadence, reachingmin_samples: 20per variant (40 total observations) takes roughly 8 weeks — within the "few months" bound for sparse-traffic eligibility, and the workflow is not dispatch-only. If the hypothesis holds, the candidate variant is expected to reduceexecution-step-count(and by extensionavg_aic, currently 186.53 for the one recorded run vs a fleet successful-run median in the 11–21 range) by bounding what is today unconstrained issue/discussion exploration in Step 3, while themodule_review_completenesseval guards against a shallower, lower-quality review.Validation
Both
./gh-aw compile go-fan --strictand./gh-aw compile go-fan --strict --validateexited0(with an expected informational⚠ Using experimental feature: graderswarning, not an error).git diff --stattouched only.github/workflows/go-fan.mdand its matching.github/workflows/go-fan.lock.yml. The compiled lock file confirmsprompt_tightening_v1is wired into the activation job'spick-experimentstep and into theGH_AW_EXPERIMENTS_PROMPT_TIGHTENING_V1env var used by the agent job templating.Interpretation
Applying this patch only starts the experiment; no decision is interpreted here. Once
min_samples: 20is reached,gh aw experiments analyze go-fancomputes the deterministicEXTEND/PROMOTE/REJECT/INCONCLUSIVEdecision unchanged — this workflow never recomputes, reinterprets, or overrides that decision, and any eventualPROMOTEstill requires a separate, human-reviewed change through the existingdaily-experiment-reportdeterministic decision engine. This workflow never merges anything itself.Rollback
git revertthe commit that applies this patch. No other files depend on this change — theexperiments:/graders:/evals:blocks and the{{#if}}conditional are fully self-contained withingo-fan.md, and reverting regenerates an unchangedgo-fan.lock.ymlviagh aw compile go-fan.Manual Patch (apply by hand)
Application Plan
proposal.patch.git apply proposal.patchgh aw compile go-fan --strict --validate