feat(loop-audit): score what is proven, not what is present - #642
Open
THRISHAL12345 wants to merge 4 commits into
Open
THRISHAL12345 wants to merge 4 commits into
THRISHAL12345 wants to merge 4 commits into
Conversation
…er it needs to load Claude Code only loads a subagent file with name and description frontmatter. This one had none, so the starter's Claude verifier was never invocable. Name and description come from its Codex twin.
Results are grouped by guardrail (gate, breaker, verifier, ...). A run replaces only the groups it drilled, so refreshing the cheap offline drills doesn't discard a slow canary. Failures and skip reasons are recorded too. The gate group carries a sha256 of the policy file, with line endings normalised so Windows and Linux checkouts agree. loop-audit uses it to tell a proof of the current gate.yaml from a proof of an older one.
A repo of 19 empty files (142 bytes) scored 100/L3, the same as the
reference repo: nearly every signal was fileExists(), and a skill counted if
its directory existed.
Placeholders no longer score. Empty or whitespace-only files, and {} / []
JSON, earn nothing and are listed under "Not counted". A skill or Claude
verifier agent needs name + description frontmatter to count. gate.yaml
needs version: 1 and a denylist. .github/ needs a non-empty file.
L3 now requires proven guardrails, read from loop-drill.json: the current
gate.yaml must pass its drills in both directions, and no recorded guardrail
may be failing. A guardrail loop-drill shows failing loses its points (gate,
verifier, breaker). Proof goes stale when gate.yaml changes, or after 30
days for canaries.
The fake repo now scores 72/L1. No starter's score changes.
Commit loop-drill.json so the reference repo stays at L3 on proof rather than files. ci-audit-gates.sh now fails when the record is missing, stale or failing, and prints the command to re-record. ci-validate-gates.sh already re-runs the drills against the live gate.yaml, so the committed record cannot claim more than they show. The L3 criteria in QUICKSTART, operating-loops and SECURITY now mention the proof.
Contributor
|
Thanks @THRISHAL12345 for contributing a docs improvement — visible, reviewable PRs like this grow the reference for everyone. What happens next
More ways to help — loop-engineering maintainers |
Contributor
|
This PR changes paths that must run the real Fork PRs from first-time contributors start with those workflows waiting for approval. A maintainer needs to open the Checks tab and click Approve and run workflows. Until that happens, branch protection will show the PR as blocked even after a review. Content-only PRs ( — loop-engineering fork-pr-gate |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The problem
loop-audit's score can be faked withtouch. This repo scores 100/100 L3:It's 19 files and 142 bytes, and it gets the same score and level as the reference repo. Almost every signal is
fileExists(), and a skill counts if its directory exists. That includes the 14-point verifier that L3 is gated on, and agate.yamlthatloop-gaterefuses to load.loop-drill(#600) can already show whether a guardrail fires, but nothing fed its results back into the score.docs/architecture-diagrams.mdsays L2 → L3 needs "denylist + budget + gates proven". The scorer never checked that last part.What this changes
1. Placeholders don't score:
tools/loop-audit{}/[], earns nothing. Each one is listed under Not counted rather than silently dropped.SKILL.mdthat hasnameanddescriptionfrontmatter, which every host needs before it will invoke a skill. The same goes for a Claude Code verifier agent. Bare directories no longer count.gate.yamlmust be a policy:version: 1and adenylist:. Anythingloop-gatewould refuse is reported as a failure..github/and workflows count only when they hold a non-empty file.This is a content check, not a quality judgement: one real line is enough. It only removes credit for files with nothing in them.
2. L3 needs proof:
loop-drill --record→loop-auditloop-drill --recordwritesloop-drill.json, with results grouped by guardrail.loop-auditreads it:gateYaml, verifier →verifier, breaker → stall detection. L3 is blockedloop-gatecouldn't drill it, sogateYamlis withdrawngate.yaml, or a canary over 30 days oldA few design points:
loop-drillalready requires. A denylist of**catches every seeded fault, so sensitivity alone can't prove a gate.gate.yamlwith line endings normalised, so Windows and Linux checkouts agree. Weakening the policy after recording drops L3 until the drills run again. Gate and breaker drills are deterministic, so they don't otherwise expire. Verifier and injection canaries do.--only verifier --recordreplaces the verifier group and keeps the gate and breaker results.loop-auditonly reads JSON, so it doesn't depend onloop-drill. That matters becauseloop-drillisn't on npm yet (see below).3. The reference repo proves its own guardrails
loop-drill.jsonis committed. The reference repo stays at 100/L3, now withGuardrails proven by loop-drill: breaker, gate.ci-audit-gates.shfails if that record is missing, stale, failing, or doesn't prove the gate, and it prints the command to re-record.ci-validate-gates.shalready re-runs the drills against the livegate.yaml. Together they mean the committed record can't claim more than the drills show, and editinggate.yamlwithout re-recording fails CI.Also
.claude/agents/verifier.mdhad no frontmatter, so Claude Code would never load it as a subagent. I gave it thenameanddescriptionfrom its Codex twin. The starter's score doesn't change, because the Codex verifier already counted.loop-auditREADME gets a Present is not proven section, and theloop-drillREADME gets Recording proof for loop-audit. Theloop-auditCHANGELOG has an Unreleased entry. The L3 rows indocs/QUICKSTART.md,docs/operating-loops.mdandSECURITY.mdare updated.Effect on scores
Before and after for the reference repo, all 17 starters and the fake repo:
loop-drill.jsonwas recordedNo real file in the repo is empty, and no real skill directory lacks a loadable
SKILL.md. I checked all 102SKILL.mdand agent files; the only one lacking frontmatter was the starter fixed above.Verification
loop-audithas 50 (20 new),loop-drillhas 69 (12 new). New coverage includes:--onlymerging.--json --recordkeeping stdout parseable.gate.yamlis weakened.sha256sum.{}JSON, frontmatter and thedescriptionrequirement, the gate shape and.githubcontent checks.ci-validate-gates.shexits 0 with 362 tests passing and none failing, including theloop-drilldogfood step.ci-audit-gates.shpasses:Reference score: 100,Reference guardrails proven: breaker, gate.before-after-demo.shgoes 7 → 64 → 94 (L2), as before.scripts/github-triage.test.mjsfails on every Windows checkout because it builds its path withURL.pathname. fix: guard loops against prompt injection from untrusted input #641 fixes that in its first commit. I applied that fix only for the local run and didn't include it here.mcp-server"server lists all tools over stdio" 10s timeout (10,083 ms) straight afternpm ci. It's the same cold-start flake noted on fix: guard loops against prompt injection from untrusted input #641. Warm, it passed in 4,553 ms.python3, which this machine doesn't have, so I ran it with apython3→pythonshim.Limits — please read
gate.yamlthat really blocks.env, keyword-stuffedLOOP.md, plus a realloop-drill --record. Its gate genuinely works, so the proof is true, but nothing judges whether the skills' instructions are any good. That limit is inherent to scoring files. The next lever would be requiring a recorded verifier canary for L3. I haven't done that, because it costs real verifier runs and would need a verifier command for this repo. That's your call.Last run:. Someone determined can hand-write one with the right sha256. It raises the bar fromtouchto deliberate forgery, and CI re-running the drills (as this repo now does) is what keeps it honest.Merge order:
loop-drillmust be on npm beforeloop-auditis released@cobusgreyling/loop-drillisn't published yet:npm viewreturns 404, and there's norelease-loop-drill.yml. Once this ships in aloop-auditrelease, L3 needs a record that users can only produce withnpx @cobusgreyling/loop-drill. So I've not bumpedloop-audit's version. The CHANGELOG entry is under Unreleased, so you can release it afterloop-drill's first publish. Merging on its own is safe: it only changes this repo's code and CI.Overlap with #641
#641 adds
--agent-cmdand aninjectiondrill toloop-drill's CLI, and this PR adds--recordto the same file. Whichever merges second will have a small, mechanical conflict incli.ts, and I'll rebase it. Grouping is by drill-id prefix, so injection results get recorded and scored with no further change. A failing injection canary then blocks L3 like any other failing guardrail.