SDD fix-loop redesign: resume-based fix rounds, five-round breaker, controller adjudication, lifecycle restructure - #1998
Merged
Merged
Conversation
arittr
added a commit
to prime-radiant-inc/superpowers-evals
that referenced
this pull request
Jul 18, 2026
Draft-only GitHub comment answering the maintainer's eval ask on obra/superpowers#1998, sourced from the final campaign log. Not posted; for Drew's review.
Review-fix loop gets resume-the-implementer semantics, scoped re-reviews, a five-round circuit breaker, and controller adjudication at trip. SKILL.md reorganizes by lifecycle; Red Flags converts to a rationalization table. Brainstormed with Jesse 2026-07-15.
Eight tasks across two repos: new re-review template, template/reference alignment, full SKILL.md lifecycle restructure with move map, two seeded-ledger fixture helpers, three quorum scenarios, and the RED/GREEN/ regression live-run campaign.
arittr
force-pushed
the
sdd-fix-loop-redesign
branch
from
July 19, 2026 19:22
1f97eda to
8136e0b
Compare
…nd breaker, and rationalization table
arittr
force-pushed
the
sdd-fix-loop-redesign
branch
from
July 19, 2026 19:30
8136e0b to
78bbfda
Compare
1 task done
lucianghinda
added a commit
to lucianghinda/superpowers-ruby
that referenced
this pull request
Aug 13, 2026
Four separable hardening fixes, each with a first-hand upstream failure report. Upstream's single-reviewer consolidation is deliberately NOT taken — its author publicly corrected their own catch-rate claim (two reviewers 4/5 against the new design's 3/5) and a user reported doubled token usage. The spec and code-quality reviewers both stay. The fix loop said "Repeat until approved" with no breaker. Reviewers are nondeterministic, so approval is not a fixed point the loop converges to — each pass samples a different subset of findings and can keep discovering new ones indefinitely. Now five rounds maximum: rounds 1-3 resume the original implementer whose context is intact, rounds 4-5 dispatch a fresh one on a more capable model, and an exhausted breaker forces the controller to adjudicate each open finding on the record — BLOCKED to the human partner if any is load-bearing. Three failed resumes is evidence about the agent rather than the task, which is why round 4 changes context and capability together. (obra#1998) That section also contradicted itself about who owns corrections: the fix loop said the same subagent fixes, the block six lines below said dispatch a fix subagent. Resolved in favour of the round schedule, with the outright-failure path scoped to BLOCKED/NEEDS_CONTEXT where it belongs. An implementer could claim TDD while giving the controller no proof the test failed first, so the report contract now requires RED and GREEN commands with their output. "I followed TDD" is not evidence. (obra#1065) The spec reviewer was told to "read the implementation code" with no bounded evidence, so reviewers re-read the repository. It now gets the task's diff range, with out-of-diff inspection allowed only against a named risk that must appear in the report. Upstream also suggested solving this by routing reviewers to a cheaper model; that model-policy assumption is not taken. (obra#1538) Reviewers attempting implementation-side repository operations caused detached-HEAD failures, so both reviewer paths are now explicitly read-only on the checkout. (obra#1543) Upstream's versions of these are entangled with its durable SDD workspace (report files, ledger, review-package script). This fork never adopted that workspace, so none of it comes along — verified absent. Ported from obra#1065, obra#1538, obra#1543 and obra#1998. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
5 tasks done
3 of 5 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Who is submitting this PR? (required)
claude-fable-5)What problem are you trying to solve?
Four problems with
subagent-driven-development, reported by the maintainer from real sessions on frontier models:| Excuse | Reality |rationalization table.What does this PR change?
Restructures SKILL.md in lifecycle order with a specified fix loop: rounds 1–3 resume the original implementer with the findings, rounds 4–5 dispatch a fresh implementer on a more capable model, and a five-round breaker triggers controller adjudication (park with a written ledger ruling, or STOP as BLOCKED when the finding is load-bearing). Re-reviews are scoped to the findings via a new
re-review-prompt.md; the implementer/task-reviewer templates and the Codex reference are aligned with resume semantics; exact ledger line formats are specified; Red Flags becomes a rationalization table. Design spec and implementation plan (with a 33-row verbatim move map) are included.Is this change appropriate for the core library?
Yes — it modifies an existing core skill only. No new dependencies, no third-party integrations, no domain-specific content.
What alternatives did you consider?
Does this PR contain multiple unrelated changes?
No. Every file serves the fix-loop redesign: the skill, its three templates, the one-sentence Codex reference alignment (implementer close timing must permit resume), and the spec/plan documents for this change.
Existing PRs
Environment tested
New harness support (required if this PR adds a new harness)
Not applicable — no new harness.
Evaluation
main(docs/experiments/2026-07-sdd-fix-loop-redesign.md, evals commit8192fe2).Rigor
Human review
@arittr — could you run evals on this? The campaign log and scenarios live in superpowers-evals
main(docs/experiments/2026-07-sdd-fix-loop-redesign.md); the three new SDD scenarios plus the four regression scenarios against this branch are the relevant set. SUPERPOWERS_ROOT should point at this branch's checkout for GREEN and adevcheckout for baseline comparison.