Summary
pr-sous-chef is the busiest tracked workflow in the 14-day window (3/3 successful runs) and already carries real instrumentation (grader:execution-duration, eval:comment-added, eval:pr-evaluated) plus one active context assembly experiment. Its own step instructions show the pr-processor sub-agent call has a hard no-retry policy on failure — this proposes a second, independent experiment testing exactly one bounded retry on that failure path, under the tool interaction dimension.
Observation
per-workflow-summary.json shows pr-sous-chef with run_count=3, success_count=3, failure_count=0, avg_aic=25.11, avg_action_minutes=10.67 — the most active and best-instrumented workflow with real 14-day traffic among all eligible candidates (most other candidates in that file are either currently broken with a single failed run, e.g. daily-grader-audit avg_aic=252/ai-moderator/daily-storify, or have only one recorded run with no active experiment infrastructure). Reading the workflow's own current instructions (.github/workflows/pr-sous-chef.md, "Step-by-Step Process" item 8) shows: "If a pr-processor call returns non-JSON or an error, record {pr_number: <N>, skip_reason: "sub_agent_error"} ... and move to the next PR without retrying." — any transient sub-agent hiccup (e.g. a single malformed/truncated model response) currently drops that PR from the run's nudge coverage permanently, with no bounded retry attempted.
Hypothesis
Mechanism: sub-agent invocations can fail transiently on a single call (e.g. a malformed JSON response from one model turn). Retrying the same pr-processor call exactly once before giving up should recover a fraction of these transient failures into successful PR evaluations/comments, increasing the comment-added pass rate. Because the retry only fires on the (rare) failure path, the common all-succeed case is unaffected, so overall execution-duration should not regress meaningfully.
Expected direction: increase in eval:comment-added pass rate.
Minimum effect: >=0.10 absolute increase in comment-added pass rate.
Candidate Mutation
- primary_dimension:
tool interaction
- subtype:
retry_policy
- control: unchanged — a failed
pr-processor call is skipped and recorded immediately, with zero retries (byte-identical to current behavior).
- candidate: the exact same failure is retried exactly once with the same PR number/context before falling back to the existing skip-and-record behavior. No other instruction, tool, or permission changes.
Experiment
retry_policy_v1:
variants: [control, candidate]
description: "Test allowing exactly one bounded retry of a failed pr-processor sub-agent call (non-JSON or error response) before falling back to the existing skip-and-record behavior, instead of skipping immediately on the first failure."
hypothesis: "H0: No meaningful difference in comment-added pass rate between control and candidate. H1: Retrying a transient pr-processor sub-agent failure once before skipping increases the comment-added pass rate without lowering pr-evaluated coverage or materially increasing execution-duration."
metric: "eval:comment-added"
guardrail_metrics:
- name: "eval:pr-evaluated"
threshold: ">=0.90"
- name: "grader:execution-duration"
threshold: "<=900000"
min_samples: 20
analysis_type: proportion_test
decision:
minimum_effect: 0.10
regression_tolerance: 0.10
confidence: 0.95
tags: ["harness_dimension:tool interaction", "harness_subtype:retry_policy"]
No new graders:/evals: block was needed — pr-sous-chef already declares graders: { execution-duration: {} } and evals: [comment-added, nudge-targeted, pr-evaluated].
Guardrails
| Guardrail |
Threshold |
Why |
eval:pr-evaluated |
>=0.90 |
Confirms the retry doesn't crowd out actually evaluating PRs (e.g. by looping instead of moving on) — protects the workflow's core coverage guarantee. |
grader:execution-duration |
<=900000 (ms, 15 min) |
Correctness/cost-equivalence guardrail: caps run duration well above the observed avg_action_minutes=10.67 (≈640s) baseline from per-workflow-summary.json, catching a runaway-retry regression while leaving headroom for the occasional extra sub-agent call on the failure path. Required because a naive retry-on-every-call implementation could otherwise silently double run cost. |
Expected Economics
pr-sous-chef runs on a 15-minute schedule and was the most active workflow observed (3 runs/14 days in per-workflow-summary.json, gated by skip-if-no-match: "is:pr is:open -is:draft -author:app/dependabot"). At that observed rate, reaching min_samples: 20 per arm (40 total observations) would take roughly 6 months of the raw 3-runs/14-days rate; however this rate reflects only runs where the prefilter found eligible PRs at all, and gh-aw is an actively developed repository with frequent open non-draft PRs, so the realized per-run "PR processed" sample count (each PR processed within a run is itself an observation for eval-based metrics) should accumulate meaningfully faster than the coarse run-count suggests. The candidate is expected to have a small positive delta on comment-added pass rate and a negligible (near-zero, bounded by the guardrail) delta on execution-duration.
Validation
./gh-aw compile pr-sous-chef --strict and ./gh-aw compile pr-sous-chef --strict --validate both exited 0 (only pre-existing, unrelated warnings: push-to-pull-request-branch target wildcard, and experimental-feature notices for graders/approve-workflow-run/gh-aw-detection). git diff --stat after compiling touched only .github/workflows/pr-sous-chef.md and .github/workflows/pr-sous-chef.lock.yml. The compiled pr-sous-chef.lock.yml contains retry_policy_v1 wired into the experiment-picker step, the GH_AW_EXPERIMENT_SPEC JSON, and per-job GH_AW_EXPERIMENTS_RETRY_POLICY_V1 env vars, confirming the block was accepted and templated correctly. The edit was verified then reverted from the working tree (this proposer has no write authority) — the exact patch is below for a human to apply.
Interpretation
Applying this patch only starts the experiment; no decision is made or implied here. Once min_samples: 20 is reached, gh aw experiments analyze pr-sous-chef computes the deterministic EXTEND/PROMOTE/REJECT/INCONCLUSIVE verdict unchanged — this workflow never recomputes, reinterprets, or overrides that decision, and any eventual PROMOTE still requires a separate, human-reviewed change through the existing daily-experiment-report deterministic decision engine. This workflow never merges anything itself.
Rollback
Revert the commit that applies this patch (git revert <commit-sha>). No other file depends on the retry_policy_v1 block or the templated retry instruction; removing it restores pr-sous-chef.md to single-retry-free behavior with no side effects on the sibling remove_redundant_context_v1 experiment.
Manual Patch (apply by hand)
diff --git a/.github/workflows/pr-sous-chef.md b/.github/workflows/pr-sous-chef.md
index f1c5348..c2df619 100644
--- a/.github/workflows/pr-sous-chef.md
+++ b/.github/workflows/pr-sous-chef.md
@@ -381,6 +381,23 @@ experiments:
regression_tolerance: 15000
confidence: 0.95
tags: ["harness_dimension:context assembly", "harness_subtype:remove_redundant_context"]
+ retry_policy_v1:
+ variants: [control, candidate]
+ description: "Test allowing exactly one bounded retry of a failed pr-processor sub-agent call (non-JSON or error response) before falling back to the existing skip-and-record behavior, instead of skipping immediately on the first failure."
+ hypothesis: "H0: No meaningful difference in comment-added pass rate between control and candidate. H1: Retrying a transient pr-processor sub-agent failure once before skipping increases the comment-added pass rate without lowering pr-evaluated coverage or materially increasing execution-duration."
+ metric: "eval:comment-added"
+ guardrail_metrics:
+ - name: "eval:pr-evaluated"
+ threshold: ">=0.90"
+ - name: "grader:execution-duration"
+ threshold: "<=900000"
+ min_samples: 20
+ analysis_type: proportion_test
+ decision:
+ minimum_effect: 0.10
+ regression_tolerance: 0.10
+ confidence: 0.95
+ tags: ["harness_dimension:tool interaction", "harness_subtype:retry_policy"]
---
# PR Sous Chef 🍳
@@ -417,7 +434,11 @@ When this workflow is triggered by the `/souschef` slash command on a PR comment
If two PRs are still tied, prioritize the lower PR number first for deterministic behavior and stable reruns.
6. After applying skip rules, stop creating new nudge comments once 4 PRs have been nudged in the current run. Continue processing only for required bookkeeping/reporting.
7. Use the `pr-processor` sub-agent for each PR; pass only the PR number and compact context.
+{{#if experiments.retry_policy_v1 == 'candidate' }}
+8. If a `pr-processor` call returns non-JSON or an error, retry the same `pr-processor` call exactly once with the same PR number and context. If the retry also returns non-JSON or an error, record `{pr_number: <N>, skip_reason: "sub_agent_error"}` in the `skipped` array of the run-summary issue payload and move to the next PR without further retries.
+{{#else}}
8. If a `pr-processor` call returns non-JSON or an error, record `{pr_number: <N>, skip_reason: "sub_agent_error"}` in the `skipped` array of the run-summary issue payload and move to the next PR without retrying.
+{{#endif}}
9. Do not fetch full PR diffs or large file lists unless absolutely required for a skip decision.
{{#if experiments.remove_redundant_context_v1 == 'control' }}
Application Plan
- Save the diff above as
proposal.patch.
git apply proposal.patch
gh aw compile pr-sous-chef --strict --validate
- Open a PR manually if compilation succeeds.
Generated by 🧫 Daily Harness Experiment Proposer · claude · agent · 207.7 AIC · ⊞ 7.1K · ◷
Summary
pr-sous-chefis the busiest tracked workflow in the 14-day window (3/3 successful runs) and already carries real instrumentation (grader:execution-duration,eval:comment-added,eval:pr-evaluated) plus one activecontext assemblyexperiment. Its own step instructions show thepr-processorsub-agent call has a hard no-retry policy on failure — this proposes a second, independent experiment testing exactly one bounded retry on that failure path, under thetool interactiondimension.Observation
per-workflow-summary.jsonshowspr-sous-chefwithrun_count=3, success_count=3, failure_count=0, avg_aic=25.11, avg_action_minutes=10.67— the most active and best-instrumented workflow with real 14-day traffic among all eligible candidates (most other candidates in that file are either currently broken with a single failed run, e.g.daily-grader-auditavg_aic=252/ai-moderator/daily-storify, or have only one recorded run with no active experiment infrastructure). Reading the workflow's own current instructions (.github/workflows/pr-sous-chef.md, "Step-by-Step Process" item 8) shows: "If apr-processorcall returns non-JSON or an error, record{pr_number: <N>, skip_reason: "sub_agent_error"}... and move to the next PR without retrying." — any transient sub-agent hiccup (e.g. a single malformed/truncated model response) currently drops that PR from the run's nudge coverage permanently, with no bounded retry attempted.Hypothesis
Mechanism: sub-agent invocations can fail transiently on a single call (e.g. a malformed JSON response from one model turn). Retrying the same
pr-processorcall exactly once before giving up should recover a fraction of these transient failures into successful PR evaluations/comments, increasing thecomment-addedpass rate. Because the retry only fires on the (rare) failure path, the common all-succeed case is unaffected, so overallexecution-durationshould not regress meaningfully.Expected direction: increase in
eval:comment-addedpass rate.Minimum effect: >=0.10 absolute increase in
comment-addedpass rate.Candidate Mutation
tool interactionretry_policypr-processorcall is skipped and recorded immediately, with zero retries (byte-identical to current behavior).Experiment
No new
graders:/evals:block was needed —pr-sous-chefalready declaresgraders: { execution-duration: {} }andevals: [comment-added, nudge-targeted, pr-evaluated].Guardrails
eval:pr-evaluated>=0.90grader:execution-duration<=900000(ms, 15 min)avg_action_minutes=10.67(≈640s) baseline fromper-workflow-summary.json, catching a runaway-retry regression while leaving headroom for the occasional extra sub-agent call on the failure path. Required because a naive retry-on-every-call implementation could otherwise silently double run cost.Expected Economics
pr-sous-chefruns on a 15-minute schedule and was the most active workflow observed (3 runs/14 days inper-workflow-summary.json, gated byskip-if-no-match: "is:pr is:open -is:draft -author:app/dependabot"). At that observed rate, reachingmin_samples: 20per arm (40 total observations) would take roughly 6 months of the raw 3-runs/14-days rate; however this rate reflects only runs where the prefilter found eligible PRs at all, andgh-awis an actively developed repository with frequent open non-draft PRs, so the realized per-run "PR processed" sample count (each PR processed within a run is itself an observation for eval-based metrics) should accumulate meaningfully faster than the coarse run-count suggests. The candidate is expected to have a small positive delta oncomment-addedpass rate and a negligible (near-zero, bounded by the guardrail) delta onexecution-duration.Validation
./gh-aw compile pr-sous-chef --strictand./gh-aw compile pr-sous-chef --strict --validateboth exited 0 (only pre-existing, unrelated warnings:push-to-pull-request-branchtarget wildcard, and experimental-feature notices forgraders/approve-workflow-run/gh-aw-detection).git diff --statafter compiling touched only.github/workflows/pr-sous-chef.mdand.github/workflows/pr-sous-chef.lock.yml. The compiledpr-sous-chef.lock.ymlcontainsretry_policy_v1wired into the experiment-picker step, theGH_AW_EXPERIMENT_SPECJSON, and per-jobGH_AW_EXPERIMENTS_RETRY_POLICY_V1env vars, confirming the block was accepted and templated correctly. The edit was verified then reverted from the working tree (this proposer has no write authority) — the exact patch is below for a human to apply.Interpretation
Applying this patch only starts the experiment; no decision is made or implied here. Once
min_samples: 20is reached,gh aw experiments analyze pr-sous-chefcomputes the deterministicEXTEND/PROMOTE/REJECT/INCONCLUSIVEverdict unchanged — this workflow never recomputes, reinterprets, or overrides that decision, and any eventualPROMOTEstill requires a separate, human-reviewed change through the existingdaily-experiment-reportdeterministic decision engine. This workflow never merges anything itself.Rollback
Revert the commit that applies this patch (
git revert <commit-sha>). No other file depends on theretry_policy_v1block or the templated retry instruction; removing it restorespr-sous-chef.mdto single-retry-free behavior with no side effects on the siblingremove_redundant_context_v1experiment.Manual Patch (apply by hand)
Application Plan
proposal.patch.git apply proposal.patchgh aw compile pr-sous-chef --strict --validate