Skip to content

[harness-experiment-proposal] pr-sous-chef — tool interaction/retry_policy A/B harness experiment #64227

Description

@github-actions

Summary

pr-sous-chef is the busiest tracked workflow in the 14-day window (3/3 successful runs) and already carries real instrumentation (grader:execution-duration, eval:comment-added, eval:pr-evaluated) plus one active context assembly experiment. Its own step instructions show the pr-processor sub-agent call has a hard no-retry policy on failure — this proposes a second, independent experiment testing exactly one bounded retry on that failure path, under the tool interaction dimension.

Observation

per-workflow-summary.json shows pr-sous-chef with run_count=3, success_count=3, failure_count=0, avg_aic=25.11, avg_action_minutes=10.67 — the most active and best-instrumented workflow with real 14-day traffic among all eligible candidates (most other candidates in that file are either currently broken with a single failed run, e.g. daily-grader-audit avg_aic=252/ai-moderator/daily-storify, or have only one recorded run with no active experiment infrastructure). Reading the workflow's own current instructions (.github/workflows/pr-sous-chef.md, "Step-by-Step Process" item 8) shows: "If a pr-processor call returns non-JSON or an error, record {pr_number: <N>, skip_reason: "sub_agent_error"} ... and move to the next PR without retrying." — any transient sub-agent hiccup (e.g. a single malformed/truncated model response) currently drops that PR from the run's nudge coverage permanently, with no bounded retry attempted.

Hypothesis

Mechanism: sub-agent invocations can fail transiently on a single call (e.g. a malformed JSON response from one model turn). Retrying the same pr-processor call exactly once before giving up should recover a fraction of these transient failures into successful PR evaluations/comments, increasing the comment-added pass rate. Because the retry only fires on the (rare) failure path, the common all-succeed case is unaffected, so overall execution-duration should not regress meaningfully.
Expected direction: increase in eval:comment-added pass rate.
Minimum effect: >=0.10 absolute increase in comment-added pass rate.

Candidate Mutation

  • primary_dimension: tool interaction
  • subtype: retry_policy
  • control: unchanged — a failed pr-processor call is skipped and recorded immediately, with zero retries (byte-identical to current behavior).
  • candidate: the exact same failure is retried exactly once with the same PR number/context before falling back to the existing skip-and-record behavior. No other instruction, tool, or permission changes.

Experiment

  retry_policy_v1:
    variants: [control, candidate]
    description: "Test allowing exactly one bounded retry of a failed pr-processor sub-agent call (non-JSON or error response) before falling back to the existing skip-and-record behavior, instead of skipping immediately on the first failure."
    hypothesis: "H0: No meaningful difference in comment-added pass rate between control and candidate. H1: Retrying a transient pr-processor sub-agent failure once before skipping increases the comment-added pass rate without lowering pr-evaluated coverage or materially increasing execution-duration."
    metric: "eval:comment-added"
    guardrail_metrics:
      - name: "eval:pr-evaluated"
        threshold: ">=0.90"
      - name: "grader:execution-duration"
        threshold: "<=900000"
    min_samples: 20
    analysis_type: proportion_test
    decision:
      minimum_effect: 0.10
      regression_tolerance: 0.10
      confidence: 0.95
    tags: ["harness_dimension:tool interaction", "harness_subtype:retry_policy"]

No new graders:/evals: block was needed — pr-sous-chef already declares graders: { execution-duration: {} } and evals: [comment-added, nudge-targeted, pr-evaluated].

Guardrails

Guardrail Threshold Why
eval:pr-evaluated >=0.90 Confirms the retry doesn't crowd out actually evaluating PRs (e.g. by looping instead of moving on) — protects the workflow's core coverage guarantee.
grader:execution-duration <=900000 (ms, 15 min) Correctness/cost-equivalence guardrail: caps run duration well above the observed avg_action_minutes=10.67 (≈640s) baseline from per-workflow-summary.json, catching a runaway-retry regression while leaving headroom for the occasional extra sub-agent call on the failure path. Required because a naive retry-on-every-call implementation could otherwise silently double run cost.

Expected Economics

pr-sous-chef runs on a 15-minute schedule and was the most active workflow observed (3 runs/14 days in per-workflow-summary.json, gated by skip-if-no-match: "is:pr is:open -is:draft -author:app/dependabot"). At that observed rate, reaching min_samples: 20 per arm (40 total observations) would take roughly 6 months of the raw 3-runs/14-days rate; however this rate reflects only runs where the prefilter found eligible PRs at all, and gh-aw is an actively developed repository with frequent open non-draft PRs, so the realized per-run "PR processed" sample count (each PR processed within a run is itself an observation for eval-based metrics) should accumulate meaningfully faster than the coarse run-count suggests. The candidate is expected to have a small positive delta on comment-added pass rate and a negligible (near-zero, bounded by the guardrail) delta on execution-duration.

Validation

./gh-aw compile pr-sous-chef --strict and ./gh-aw compile pr-sous-chef --strict --validate both exited 0 (only pre-existing, unrelated warnings: push-to-pull-request-branch target wildcard, and experimental-feature notices for graders/approve-workflow-run/gh-aw-detection). git diff --stat after compiling touched only .github/workflows/pr-sous-chef.md and .github/workflows/pr-sous-chef.lock.yml. The compiled pr-sous-chef.lock.yml contains retry_policy_v1 wired into the experiment-picker step, the GH_AW_EXPERIMENT_SPEC JSON, and per-job GH_AW_EXPERIMENTS_RETRY_POLICY_V1 env vars, confirming the block was accepted and templated correctly. The edit was verified then reverted from the working tree (this proposer has no write authority) — the exact patch is below for a human to apply.

Interpretation

Applying this patch only starts the experiment; no decision is made or implied here. Once min_samples: 20 is reached, gh aw experiments analyze pr-sous-chef computes the deterministic EXTEND/PROMOTE/REJECT/INCONCLUSIVE verdict unchanged — this workflow never recomputes, reinterprets, or overrides that decision, and any eventual PROMOTE still requires a separate, human-reviewed change through the existing daily-experiment-report deterministic decision engine. This workflow never merges anything itself.

Rollback

Revert the commit that applies this patch (git revert <commit-sha>). No other file depends on the retry_policy_v1 block or the templated retry instruction; removing it restores pr-sous-chef.md to single-retry-free behavior with no side effects on the sibling remove_redundant_context_v1 experiment.

Manual Patch (apply by hand)

diff --git a/.github/workflows/pr-sous-chef.md b/.github/workflows/pr-sous-chef.md
index f1c5348..c2df619 100644
--- a/.github/workflows/pr-sous-chef.md
+++ b/.github/workflows/pr-sous-chef.md
@@ -381,6 +381,23 @@ experiments:
       regression_tolerance: 15000
       confidence: 0.95
     tags: ["harness_dimension:context assembly", "harness_subtype:remove_redundant_context"]
+  retry_policy_v1:
+    variants: [control, candidate]
+    description: "Test allowing exactly one bounded retry of a failed pr-processor sub-agent call (non-JSON or error response) before falling back to the existing skip-and-record behavior, instead of skipping immediately on the first failure."
+    hypothesis: "H0: No meaningful difference in comment-added pass rate between control and candidate. H1: Retrying a transient pr-processor sub-agent failure once before skipping increases the comment-added pass rate without lowering pr-evaluated coverage or materially increasing execution-duration."
+    metric: "eval:comment-added"
+    guardrail_metrics:
+      - name: "eval:pr-evaluated"
+        threshold: ">=0.90"
+      - name: "grader:execution-duration"
+        threshold: "<=900000"
+    min_samples: 20
+    analysis_type: proportion_test
+    decision:
+      minimum_effect: 0.10
+      regression_tolerance: 0.10
+      confidence: 0.95
+    tags: ["harness_dimension:tool interaction", "harness_subtype:retry_policy"]
 ---
 
 # PR Sous Chef 🍳
@@ -417,7 +434,11 @@ When this workflow is triggered by the `/souschef` slash command on a PR comment
    If two PRs are still tied, prioritize the lower PR number first for deterministic behavior and stable reruns.
 6. After applying skip rules, stop creating new nudge comments once 4 PRs have been nudged in the current run. Continue processing only for required bookkeeping/reporting.
 7. Use the `pr-processor` sub-agent for each PR; pass only the PR number and compact context.
+{{#if experiments.retry_policy_v1 == 'candidate' }}
+8. If a `pr-processor` call returns non-JSON or an error, retry the same `pr-processor` call exactly once with the same PR number and context. If the retry also returns non-JSON or an error, record `{pr_number: <N>, skip_reason: "sub_agent_error"}` in the `skipped` array of the run-summary issue payload and move to the next PR without further retries.
+{{#else}}
 8. If a `pr-processor` call returns non-JSON or an error, record `{pr_number: <N>, skip_reason: "sub_agent_error"}` in the `skipped` array of the run-summary issue payload and move to the next PR without retrying.
+{{#endif}}
 9. Do not fetch full PR diffs or large file lists unless absolutely required for a skip decision.
 
 {{#if experiments.remove_redundant_context_v1 == 'control' }}

Application Plan

  1. Save the diff above as proposal.patch.
  2. git apply proposal.patch
  3. gh aw compile pr-sous-chef --strict --validate
  4. Open a PR manually if compilation succeeds.

Generated by 🧫 Daily Harness Experiment Proposer · claude · agent · 207.7 AIC · ⊞ 7.1K · ◷

  • expires on Oct 6, 2026, 12:49 AM UTC-08:00

Activity

  1. github-actions commented on Oct 6, 2026

    @github-actions
    ContributorAuthor

    This issue was automatically closed because it expired on 2026-10-06T08:49:35.744Z.

    Closed by Workflow

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions