Skip to content

feat(selector): literature-driven disagreement router v2 - #6

Merged
LeoLin990405 merged 2 commits into
mainfrom
feat/selector-v2-literature
Jul 2, 2026
Merged

LeoLin990405 merged 2 commits into
mainfrom
feat/selector-v2-literature

Conversation

@LeoLin990405

Copy link
Copy Markdown
Collaborator

What

Selector v2 — a verifier-aware disagreement router for best-of-N fan-out, grounded in 6 papers from the local reading notes and validated on the megabench 60-task trap benchmark.

Mechanisms (each ← a paper)

  • Three-way outcome TRUST / TRUST_SPOT_CHECK / ESCALATE — EDV (self-confirmation trap): unverified consensus is never clean trust.
  • Forced-escalate categories (security/correctness/impossible) — SkillHarness safe-skill boundaries + B3 data showing consensus fails correlated there.
  • Laplace-smoothed confidence (k+1)/(n+2) — OmniOPD Bayesian smoothing: a 5/5 fleet reads as 0.857, not certainty.
  • escalationPriority — OmniOPD peak-entropy scheduling (discrete): spend premium budget on the most-split tasks first.
  • Escalation as first-class action — Agentic Abstention.
  • skeptic.md playbook — Agentic Abstention CONVOLVE: B3 collective-failure traps distilled into category-level challenge rules (with provenance).

Empirical (offline replay on 60-task benchmark, strict dual-judge)

Config avoid/60 premium calls
v1 binary 44 0%
v2 (3-way + forced + smoothed) 52 43%
v2 + skeptic 58 30%
v2 + skeptic + synthetic gate 58 22%

Approaches the best premium single (59/60) at a fraction of premium cost. ⚠️ skeptic numbers are in-sample (playbook distilled from these traps) — held-out validation is the next step (deliberately out of scope here).

Test

Pure domain module. 18 vitest + fast-check cases; full suite 680 green; check:docs unaffected (domain tests not gated). Codex-reviewed ACCEPTED.

Not in this PR

Runtime wiring into dispatch/loop, playbook auto-evolution (SelfHarnessLoop), learned TRINITY-style router — all follow-ups.

Three-way outcome (TRUST/TRUST_SPOT_CHECK/ESCALATE), forced-escalate
categories, Laplace-smoothed confidence, and escalationPriority — each
mechanism grounded in a read paper (EDV, Agentic Abstention, OmniOPD,
SkillHarness). Adds a CONVOLVE-style skeptic playbook template distilled
from the megabench B3 collective-failure traps.

Pure domain module; 18 vitest+fast-check cases; full suite 680 green;
check:docs unaffected. Codex-reviewed ACCEPTED.
@LeoLin990405

Copy link
Copy Markdown
Collaborator Author

Held-out validation (follow-up to the in-sample caveat)

The PR body flagged the skeptic-playbook numbers as in-sample (playbook distilled from the same 60 traps). Ran the held-out check: 20 brand-new trap questions, same 10 categories, different specific traps the playbook never saw.

condition weak-fleet trap-avoidance
baseline (no playbook) 59/100 (59%)
+ skeptic playbook 84/100 (84%)
held-out gain +25pp
  • All 6 baseline collective-blind-spots flipped to majority-catch; 0 regressions (no over-abstention).
  • Gains spread across 9/10 categories and all 5 weak models (+3…+9 each) — not concentrated.
  • Strict dual-judge, agreement 82%.

Conclusion: the CONVOLVE-style category-level rules generalize to unseen traps (they're not memorized questions). The in-sample 58/60 overstates, but the transferable component is real and large. Data/scripts archived alongside the megabench reports.

@LeoLin990405
LeoLin990405 merged commit faad733 into main Jul 2, 2026
4 checks passed
@LeoLin990405
LeoLin990405 deleted the feat/selector-v2-literature branch July 2, 2026 10:46
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant