Skip to content

SDD fix-loop redesign: resume-based fix rounds, five-round breaker, controller adjudication, lifecycle restructure - #1998

Merged
arittr merged 5 commits into
devfrom
sdd-fix-loop-redesign
Jul 19, 2026
Merged

arittr merged 5 commits into
devfrom
sdd-fix-loop-redesign

Conversation

@obra

@obra obra commented Jul 17, 2026

Copy link
Copy Markdown
Owner

Who is submitting this PR? (required)

Field Value
Your model + version Claude Fable 5 (claude-fable-5)
Harness + version Claude Code CLI 2.1.211 (macOS)
All plugins installed superpowers 6.1.1, superpowers-chrome, elements-of-style, claude-session-driver, frontend-design, code-simplifier, dataviz, agent-sdk-dev, claude-code-setup, claude-md-management, context7, gopls-lsp, linear, mcp-server-dev, playground, plus personal skills (roborev, homedir/obsidian/macos tooling)
Human partner who reviewed this diff Jesse Vincent (@obra) — maintainer-directed session; design, plan, and every eval-campaign decision approved interactively

What problem are you trying to solve?

Four problems with subagent-driven-development, reported by the maintainer from real sessions on frontier models:

  1. Pathological review loops with no circuit breaker — implement, review, fix, review, review, fix — because the loop was literally "Repeat until approved" and every re-review was a fresh full review of the whole diff, so a nondeterministic reviewer surfaced new findings each round.
  2. A three-way contradiction about who fixes findings: the process diagram and dispatch guidance said "dispatch fix subagents"; the Red Flags section said "Implementer (same subagent) fixes them"; implementer-prompt.md assumed re-engagement. Live baseline evidence below shows this produced coin-flip behavior.
  3. Accreted structure: 13 top-level sections with guidance for one activity scattered across four of them ("Constructing Reviewer Prompts" held reviewer guidance, fix policy, final-review policy, and plan-conflict adjudication).
  4. Red Flags bullet list where the other seven skills use the | Excuse | Reality | rationalization table.

What does this PR change?

Restructures SKILL.md in lifecycle order with a specified fix loop: rounds 1–3 resume the original implementer with the findings, rounds 4–5 dispatch a fresh implementer on a more capable model, and a five-round breaker triggers controller adjudication (park with a written ledger ruling, or STOP as BLOCKED when the finding is load-bearing). Re-reviews are scoped to the findings via a new re-review-prompt.md; the implementer/task-reviewer templates and the Codex reference are aligned with resume semantics; exact ledger line formats are specified; Red Flags becomes a rationalization table. Design spec and implementation plan (with a 33-row verbatim move map) are included.

Is this change appropriate for the core library?

Yes — it modifies an existing core skill only. No new dependencies, no third-party integrations, no domain-specific content.

What alternatives did you consider?

  • Fresh fix subagents with full task context (portable, fresh eyes) and keeping dedicated narrow fixers (cheapest) — rejected for resume-the-implementer because the implementer already holds task context and ownership; narrow fixers rebuilding context per finding is plausibly part of why loops didn't converge. Harnesses without agent resume get a specified fallback (fresh dispatch carrying brief + report file + findings).
  • Full re-reviews with only a hard cap — rejected: rounds stay nondeterministic and you hit the cap with open findings more often. Scoped re-reviews make the loop structurally convergent; the final whole-branch review remains the broad safety net.
  • Human checkpoint at the breaker — rejected by the maintainer: SDD's point is autonomous execution. The breaker routes churn to the ledger and genuine plan defects to the BLOCKED stop that has always existed.
  • Targeted surgery instead of lifecycle restructure — rejected: the mixed-concerns complaint was the root issue; the restructure moves eval-tuned sentences verbatim (enforced by the move map) rather than rewording them.

Does this PR contain multiple unrelated changes?

No. Every file serves the fix-loop redesign: the skill, its three templates, the one-sentence Codex reference alignment (implementer close timing must permit resume), and the spec/plan documents for this change.

Existing PRs

Environment tested

Harness (e.g. Claude Code, Cursor) Harness version Model Model version/ID
Claude Code (quorum eval container, Linux) 2.1.209 Claude Opus 4.8 claude-opus-4-8
Claude Code (macOS, authoring + dogfooding session) 2.1.211 Claude Fable 5 claude-fable-5

New harness support (required if this PR adds a new harness)

Not applicable — no new harness.

Evaluation

  • Initial prompt: the maintainer opened the session describing the four problems above ("really pathological 'implement, review, fix, review, review, fix' loops with no circuit breaker", the fix-subagent switch, the missing rationalization table, the structural mess).
  • Eval sessions after the change: 7 live quorum runs against this branch (3 new scenarios GREEN + 4 regression), plus 5 baseline (RED) runs against dev, in the superpowers-evals lab. Scenarios, fixtures, and the full campaign log landed on evals main (docs/experiments/2026-07-sdd-fix-loop-redesign.md, evals commit 8192fe2).
  • Before/after:
    • Structural finding at the breaker: dev silently downgraded an Important plan-contradiction to Minor and built the dependent task on top of it; this branch stopped and surfaced it as BLOCKED.
    • Cap adjudication: dev left no durable record (no parked/ruling lines); this branch parked the finding with a written ruling, dispatched no round 6, completed the remaining task, and informed the final review — 8/8 deterministic checks.
    • Fix mechanism: dev took opposite mechanisms in two runs of the same scenario (dedicated fix dispatch — including one told to "ignore the trailing-newline requirement", i.e. controller pre-judging — versus same-subagent resume), demonstrating the contradiction; this branch resumed the original implementer with a scoped re-review, including organically (SendMessage ×2) in the planted-defect regression scenario.
    • Regressions: planted-defect catching, YAGNI rejection, broken-plan escalation, and spec-constraint preservation all pass on this branch.
    • Negative results are documented at equal billing in the experiment log (three defeated scenario iterations — the redesigned pre-flight legitimately defuses plan-visible seeds — plus judge stalls and a legacy-check brittleness fix).

Rigor

  • If this is a skills change: skill-behavior evaluation was run via the superpowers-evals quorum lab (RED/GREEN/regression campaign above) rather than ad-hoc pressure prompts; adversarial verification came from independent per-task reviews, two whole-branch reviews (one on the most capable model, which mechanically verified all 111 old-content chunks against the 33-row move map), and a final fix wave with scoped re-review
  • This change was tested adversarially, not just on the happy path (baseline RED runs, seeded mid-loop fixtures, a never-satisfied-reviewer state, structural-contradiction fixture)
  • I did not modify carefully-tuned content without evidence: eval-tuned sentences moved verbatim per the move map; the only licensed rewordings are fix-policy sentences, each enumerated in the plan. One known exception for maintainer waiver: the tail of old Never-item 11 ("the plan's example code is a starting point, not evidence that its weaknesses were chosen") survives in content (§3 pre-judging bullet + task-reviewer template) but not as a verbatim sentence.

Human review

  • A human has reviewed the COMPLETE proposed diff before submission

@arittr — could you run evals on this? The campaign log and scenarios live in superpowers-evals main (docs/experiments/2026-07-sdd-fix-loop-redesign.md); the three new SDD scenarios plus the four regression scenarios against this branch are the relevant set. SUPERPOWERS_ROOT should point at this branch's checkout for GREEN and a dev checkout for baseline comparison.

arittr added a commit to prime-radiant-inc/superpowers-evals that referenced this pull request Jul 18, 2026
Draft-only GitHub comment answering the maintainer's eval ask on
obra/superpowers#1998, sourced from the final campaign log. Not posted;
for Drew's review.
obra added 4 commits July 19, 2026 12:07
Review-fix loop gets resume-the-implementer semantics, scoped
re-reviews, a five-round circuit breaker, and controller adjudication
at trip. SKILL.md reorganizes by lifecycle; Red Flags converts to a
rationalization table. Brainstormed with Jesse 2026-07-15.
Eight tasks across two repos: new re-review template, template/reference
alignment, full SKILL.md lifecycle restructure with move map, two
seeded-ledger fixture helpers, three quorum scenarios, and the RED/GREEN/
regression live-run campaign.
@arittr
arittr force-pushed the sdd-fix-loop-redesign branch from 1f97eda to 8136e0b Compare July 19, 2026 19:22
@arittr
arittr force-pushed the sdd-fix-loop-redesign branch from 8136e0b to 78bbfda Compare July 19, 2026 19:30

@arittr arittr left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@arittr
arittr merged commit cc69047 into dev Jul 19, 2026
@arittr
arittr deleted the sdd-fix-loop-redesign branch July 20, 2026 23:38
lucianghinda added a commit to lucianghinda/superpowers-ruby that referenced this pull request Aug 13, 2026
Four separable hardening fixes, each with a first-hand upstream failure
report. Upstream's single-reviewer consolidation is deliberately NOT taken —
its author publicly corrected their own catch-rate claim (two reviewers 4/5
against the new design's 3/5) and a user reported doubled token usage. The
spec and code-quality reviewers both stay.

The fix loop said "Repeat until approved" with no breaker. Reviewers are
nondeterministic, so approval is not a fixed point the loop converges to —
each pass samples a different subset of findings and can keep discovering
new ones indefinitely. Now five rounds maximum: rounds 1-3 resume the
original implementer whose context is intact, rounds 4-5 dispatch a fresh
one on a more capable model, and an exhausted breaker forces the controller
to adjudicate each open finding on the record — BLOCKED to the human partner
if any is load-bearing. Three failed resumes is evidence about the agent
rather than the task, which is why round 4 changes context and capability
together. (obra#1998)

That section also contradicted itself about who owns corrections: the fix
loop said the same subagent fixes, the block six lines below said dispatch a
fix subagent. Resolved in favour of the round schedule, with the
outright-failure path scoped to BLOCKED/NEEDS_CONTEXT where it belongs.

An implementer could claim TDD while giving the controller no proof the test
failed first, so the report contract now requires RED and GREEN commands
with their output. "I followed TDD" is not evidence. (obra#1065)

The spec reviewer was told to "read the implementation code" with no bounded
evidence, so reviewers re-read the repository. It now gets the task's diff
range, with out-of-diff inspection allowed only against a named risk that
must appear in the report. Upstream also suggested solving this by routing
reviewers to a cheaper model; that model-policy assumption is not taken.
(obra#1538)

Reviewers attempting implementation-side repository operations caused
detached-HEAD failures, so both reviewer paths are now explicitly read-only
on the checkout. (obra#1543)

Upstream's versions of these are entangled with its durable SDD workspace
(report files, ledger, review-package script). This fork never adopted that
workspace, so none of it comes along — verified absent.

Ported from obra#1065, obra#1538, obra#1543 and obra#1998.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants