Skip to content

[harness-experiment-proposal] go-fan — generation control/prompt_tightening A/B harness experiment #66840

Description

@github-actions

Summary

go-fan (Go Fan — Daily Go Module Reviewer) is the only workflow in the current evidence window that succeeded yet logged an avg_aic an order of magnitude above every other successful run. This proposes a generation control/prompt_tightening experiment that bounds its open-ended Step 3 research scope, aiming to cut agent turns (execution-step-count) without degrading review completeness.

Observation

In per-workflow-summary.json, Go Fan's one recorded run succeeded (failure_count=0) but logged avg_aic=186.528795 — 6–17x higher than every other successful run in the same slice (daily-experiment-report=21.27, Daily Syntax Error Quality=13.06, Issue Monster=11.19, aw-failure-investigator=0). Reading .github/workflows/go-fan.md, Step 3 ("Research the Module") instructs the agent to open-endedly "explore the module's repository" — README, releases/changelog, and "popular usage examples in issues/discussions" — with no bound on how many tool calls/turns that exploration may take.

Hypothesis

Mechanism: Bounding Step 3 to a fixed, prioritized set of research actions (README quickstart section, latest 3 releases/changelog entries, up to 5 targeted usage-pattern search results) removes the unconstrained issue/discussion browsing that currently has no stopping condition, cutting the number of LLM turns spent on research before the agent moves to Step 4 (Serena analysis) and beyond.
Expected direction: decrease in execution-step-count (total LLM request count).
Minimum effect: >=3 absolute count decrease in execution-step-count, averaged per run.

Candidate Mutation

  • primary_dimension: generation control
  • subtype: prompt_tightening
  • Control: Step 3 keeps its existing open-ended instructions — explore the repository's README, releases/changelog, and issue/discussion usage examples with no bound — unchanged from today.
  • Candidate: Step 3's "3.1 GitHub Repository" and "3.3 Recent Updates" subsections are replaced with a bounded version: fetch only the README's primary usage section, the latest 3 releases/changelog entries, and up to 5 targeted usage-pattern search results, explicitly skipping further issue/discussion browsing. "3.2 Documentation" is untouched in both branches.

Experiment

graders:
  execution-step-count: {}
  tool-failure-count: {}

evals:
  - id: module_review_completeness
    question: "Did the agent produce a complete module review — covering the module's overview, current project usage, research findings, and at least one improvement opportunity — saved to scratchpad/mods/<module>.md, plus a matching issue?"

experiments:
  prompt_tightening_v1:
    variants: [control, candidate]
    description: "Bound Step 3 (Research the Module) to a fixed, prioritized set of research actions (README quickstart, latest 3 releases/changelog entries, up to 5 targeted usage-pattern search results) instead of open-ended repository/issue/discussion exploration, to cut the number of agent turns spent on unconstrained research."
    hypothesis: "H0: No meaningful difference in execution-step-count between control and candidate. H1: Bounding Step 3's research scope decreases execution-step-count without lowering the module_review_completeness eval pass rate."
    metric: "grader:execution-step-count"
    guardrail_metrics:
      - name: "eval:module_review_completeness"
        threshold: ">=0.90"
      - name: "grader:tool-failure-count"
        threshold: "<=5"
    min_samples: 20
    analysis_type: mann_whitney
    decision:
      minimum_effect: 3
      regression_tolerance: 3
      confidence: 0.95
    tags: ["harness_dimension:generation control", "harness_subtype:prompt_tightening"]

Guardrails

Guardrail Threshold Why it protects against regression
eval:module_review_completeness >=0.90 Since execution-step-count (a cost/efficiency proxy) is the primary metric, this eval verifies the bounded research scope still yields a complete module review (overview, usage, findings, >=1 improvement opportunity, issue + scratchpad summary) — protecting correctness/completeness equivalence against the candidate.
grader:tool-failure-count <=5 Guards against the tightened instructions causing new tool-call confusion or failures relative to today's baseline behavior.

Expected Economics

Go Fan runs on a cron: daily around 7:00 on weekdays schedule (~5 runs/week). At that cadence, reaching min_samples: 20 per variant (40 total observations) takes roughly 8 weeks — within the "few months" bound for sparse-traffic eligibility, and the workflow is not dispatch-only. If the hypothesis holds, the candidate variant is expected to reduce execution-step-count (and by extension avg_aic, currently 186.53 for the one recorded run vs a fleet successful-run median in the 11–21 range) by bounding what is today unconstrained issue/discussion exploration in Step 3, while the module_review_completeness eval guards against a shallower, lower-quality review.

Validation

Both ./gh-aw compile go-fan --strict and ./gh-aw compile go-fan --strict --validate exited 0 (with an expected informational ⚠ Using experimental feature: graders warning, not an error). git diff --stat touched only .github/workflows/go-fan.md and its matching .github/workflows/go-fan.lock.yml. The compiled lock file confirms prompt_tightening_v1 is wired into the activation job's pick-experiment step and into the GH_AW_EXPERIMENTS_PROMPT_TIGHTENING_V1 env var used by the agent job templating.

Interpretation

Applying this patch only starts the experiment; no decision is interpreted here. Once min_samples: 20 is reached, gh aw experiments analyze go-fan computes the deterministic EXTEND/PROMOTE/REJECT/INCONCLUSIVE decision unchanged — this workflow never recomputes, reinterprets, or overrides that decision, and any eventual PROMOTE still requires a separate, human-reviewed change through the existing daily-experiment-report deterministic decision engine. This workflow never merges anything itself.

Rollback

git revert the commit that applies this patch. No other files depend on this change — the experiments:/graders:/evals: blocks and the {{#if}} conditional are fully self-contained within go-fan.md, and reverting regenerates an unchanged go-fan.lock.yml via gh aw compile go-fan.

Manual Patch (apply by hand)

diff --git a/.github/workflows/go-fan.md b/.github/workflows/go-fan.md
index f515da9..a921642 100644
--- a/.github/workflows/go-fan.md
+++ b/.github/workflows/go-fan.md
@@ -31,6 +31,34 @@ engine: claude
 name: Go Fan
 strict: true
 timeout-minutes: 30
+
+graders:
+  execution-step-count: {}
+  tool-failure-count: {}
+
+evals:
+  - id: module_review_completeness
+    question: "Did the agent produce a complete module review — covering the module's overview, current project usage, research findings, and at least one improvement opportunity — saved to scratchpad/mods/<module>.md, plus a matching issue?"
+
+experiments:
+  prompt_tightening_v1:
+    variants: [control, candidate]
+    description: "Bound Step 3 (Research the Module) to a fixed, prioritized set of research actions (README quickstart, latest 3 releases/changelog entries, up to 5 targeted usage-pattern search results) instead of open-ended repository/issue/discussion exploration, to cut the number of agent turns spent on unconstrained research."
+    hypothesis: "H0: No meaningful difference in execution-step-count between control and candidate. H1: Bounding Step 3's research scope decreases execution-step-count without lowering the module_review_completeness eval pass rate."
+    metric: "grader:execution-step-count"
+    guardrail_metrics:
+      - name: "eval:module_review_completeness"
+        threshold: ">=0.90"
+      - name: "grader:tool-failure-count"
+        threshold: "<=5"
+    min_samples: 20
+    analysis_type: mann_whitney
+    decision:
+      minimum_effect: 3
+      regression_tolerance: 3
+      confidence: 0.95
+    tags: ["harness_dimension:generation control", "harness_subtype:prompt_tightening"]
+
 tools:
   bash:
   - cat go.mod
@@ -131,6 +159,29 @@ From the sorted list (most recent first):
 
 ## Step 3: Research the Module
 
+{{#if experiments.prompt_tightening_v1 == 'candidate' }}
+For the selected module, research it with a **bounded, prioritized** scope — do not open-endedly browse the repository:
+
+### 3.1 GitHub Repository (bounded)
+Use GitHub tools to fetch only:
+- The README's primary usage/quickstart section
+- The latest 3 releases (or changelog entries if no releases)
+- Up to 5 search results for usage examples in issues/discussions — skip further issue/discussion browsing beyond this
+
+### 3.2 Documentation
+Note key features and API patterns:
+- Core APIs and their purposes
+- Common usage patterns
+- Performance considerations
+- Recommended configurations
+
+### 3.3 Recent Updates
+From the releases/changelog already fetched in 3.1, note:
+- New features in recent releases
+- Breaking changes
+- Deprecations
+- Security advisories
+{{#else}}
 For the selected module, research its:
 
 ### 3.1 GitHub Repository
@@ -153,6 +204,7 @@ Check for:
 - Breaking changes
 - Deprecations
 - Security advisories
+{{#endif}}
 
 ## Step 4: Analyze Project Usage with Serena
 

Application Plan

  1. Save the diff above as proposal.patch.
  2. git apply proposal.patch
  3. gh aw compile go-fan --strict --validate
  4. Open a PR manually if compilation succeeds.

Generated by 🧫 Daily Harness Experiment Proposer · claude · agent · 203.7 AIC · ⊞ 7.3K · ◷

  • expires on Oct 15, 2026, 12:51 AM UTC-08:00
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions