Description
daily-harness-experiment-proposer.md (lines 275-276 and 310) instructs the proposer to add a workflow-specific operational-value grader as an individual, one-off PR per target workflow whenever a candidate workflow lacks a deterministic primary_metric. Prompt Clustering Analysis #64995 (2026-10-02) measured the resulting PR cluster over the last 30 days: 60 PRs, 0% merge rate — 90% (54/60) closed with zero review/comments, median open-to-close time of 15 minutes, and bulk-closed under the pr-batch:operational-value-grading / pr-action:batch_review labels that pr-triage-agent.md (lines 109, 123) applies to its low/medium-risk batch-review bucket. Even the one PR that got genuine review (#60635, 12 review comments on grader conclusion-job wiring) was still closed, not merged.
This isn't individual PRs being rejected on their merits — it's a structural mismatch: the proposer generates these as discrete tasks, but the actual review process treats them as a batch that gets triaged and closed together rather than merged one at a time. 10% of all copilot-agent PR volume (60/589 over the period) is being spent on work that structurally cannot land.
Expected Impact
Stops ~60 PRs/month of wasted agent + reviewer effort on a pattern with a confirmed 0% merge rate. Frees that task-generation capacity for the proposer's other work (experiment proposals on workflows that already have a usable metric), and/or redirects operational-value-grader additions through a batch/contract process (per this repo's own operational-value-designer skill and the AGENTS.md "Operational-value contract" rule) instead of one-off generation.
Suggested Fix
In .github/workflows/daily-harness-experiment-proposer.md, change the Phase 3/4 guidance (around lines 275-276 and 310) so that when a target workflow needs a new graders.operational-value entry, the proposer either:
- batches multiple candidate workflows' grader additions into a single PR per cycle (matching how
pr-triage-agent.md already expects to review this category in bulk), or
- defers grader creation to a separate, explicitly batch-oriented process, and only proposes the
experiments: block once a workflow already has (or itself adds, via option 1/3 in the metric-priority policy) a usable grader:<id>/eval:<id>.
Suggested Agent
Existing agent — this is a scope/sequencing change to daily-harness-experiment-proposer.md itself (self-serve fix for its own task-generation pattern).
Estimated Effort
Medium (1-4 hours) — requires careful editing of the proposer's Phase 3/4 metric-priority instructions without breaking its existing non-grader experiment-proposal flow, plus a test run to confirm it no longer emits one-off grader PRs.
Data Source
DeepReport Intelligence Briefing — 2026-10-02 (~12:55Z cycle), sourced from Prompt Clustering Analysis #64995 and cross-referenced against pr-triage-agent.md and daily-harness-experiment-proposer.md in the live repo.
Generated by 🔬 Deep Report · claude · agent · 187.4 AIC · ⌖ 7.73 AIC · ⊞ 7.3K · ◷
Description
daily-harness-experiment-proposer.md(lines 275-276 and 310) instructs the proposer to add a workflow-specificoperational-valuegrader as an individual, one-off PR per target workflow whenever a candidate workflow lacks a deterministicprimary_metric. Prompt Clustering Analysis #64995 (2026-10-02) measured the resulting PR cluster over the last 30 days: 60 PRs, 0% merge rate — 90% (54/60) closed with zero review/comments, median open-to-close time of 15 minutes, and bulk-closed under thepr-batch:operational-value-grading/pr-action:batch_reviewlabels thatpr-triage-agent.md(lines 109, 123) applies to its low/medium-risk batch-review bucket. Even the one PR that got genuine review (#60635, 12 review comments on grader conclusion-job wiring) was still closed, not merged.This isn't individual PRs being rejected on their merits — it's a structural mismatch: the proposer generates these as discrete tasks, but the actual review process treats them as a batch that gets triaged and closed together rather than merged one at a time. 10% of all copilot-agent PR volume (60/589 over the period) is being spent on work that structurally cannot land.
Expected Impact
Stops ~60 PRs/month of wasted agent + reviewer effort on a pattern with a confirmed 0% merge rate. Frees that task-generation capacity for the proposer's other work (experiment proposals on workflows that already have a usable metric), and/or redirects operational-value-grader additions through a batch/contract process (per this repo's own
operational-value-designerskill and theAGENTS.md"Operational-value contract" rule) instead of one-off generation.Suggested Fix
In
.github/workflows/daily-harness-experiment-proposer.md, change the Phase 3/4 guidance (around lines 275-276 and 310) so that when a target workflow needs a newgraders.operational-valueentry, the proposer either:pr-triage-agent.mdalready expects to review this category in bulk), orexperiments:block once a workflow already has (or itself adds, via option 1/3 in the metric-priority policy) a usablegrader:<id>/eval:<id>.Suggested Agent
Existing agent — this is a scope/sequencing change to
daily-harness-experiment-proposer.mditself (self-serve fix for its own task-generation pattern).Estimated Effort
Medium (1-4 hours) — requires careful editing of the proposer's Phase 3/4 metric-priority instructions without breaking its existing non-grader experiment-proposal flow, plus a test run to confirm it no longer emits one-off grader PRs.
Data Source
DeepReport Intelligence Briefing — 2026-10-02 (~12:55Z cycle), sourced from Prompt Clustering Analysis #64995 and cross-referenced against
pr-triage-agent.mdanddaily-harness-experiment-proposer.mdin the live repo.