Skip to content

feat(loop-audit): score what is proven, not what is present - #642

Open
THRISHAL12345 wants to merge 4 commits into
cobusgreyling:mainfrom
THRISHAL12345:feat/audit-proven
Open

THRISHAL12345 wants to merge 4 commits into
cobusgreyling:mainfrom
THRISHAL12345:feat/audit-proven

Conversation

@THRISHAL12345

Copy link
Copy Markdown
Contributor

The problem

loop-audit's score can be faked with touch. This repo scores 100/100 L3:

printf "Last run: $(date -u +%F)\n" > STATE.md
printf "gate denylist safety budget kill switch worktree MCP escalate stuck circuit breaker allowlist\n" > LOOP.md
for s in loop-triage loop-verifier minimal-fix loop-constraints loop-budget; do mkdir -p skills/$s; : > skills/$s/SKILL.md; done
mkdir -p docs .github/workflows; : > docs/safety.md; : > .github/workflows/x.yml
for f in AGENTS.md gate.yaml loop-budget.md loop-constraints.md memory-tiers.md memory-budget.md fleet-registry.md fleet-inbox.md; do : > $f; done
echo '{}' > .mcp.json; echo "{\"run_id\":\"$(date -u +%F)\"}" > loop-run-log.md

It's 19 files and 142 bytes, and it gets the same score and level as the reference repo. Almost every signal is fileExists(), and a skill counts if its directory exists. That includes the 14-point verifier that L3 is gated on, and a gate.yaml that loop-gate refuses to load.

loop-drill (#600) can already show whether a guardrail fires, but nothing fed its results back into the score. docs/architecture-diagrams.md says L2 → L3 needs "denylist + budget + gates proven". The scorer never checked that last part.

What this changes

1. Placeholders don't score: tools/loop-audit

  • Empty files are placeholders. A file that is empty or whitespace-only, or JSON that is just {}/[], earns nothing. Each one is listed under Not counted rather than silently dropped.
  • Skills must load. A skill is a directory with a SKILL.md that has name and description frontmatter, which every host needs before it will invoke a skill. The same goes for a Claude Code verifier agent. Bare directories no longer count.
  • gate.yaml must be a policy: version: 1 and a denylist:. Anything loop-gate would refuse is reported as a failure.
  • .github/ and workflows count only when they hold a non-empty file.

This is a content check, not a quality judgement: one real line is enough. It only removes credit for files with nothing in them.

2. L3 needs proof: loop-drill --record → loop-audit

loop-drill --record writes loop-drill.json, with results grouped by guardrail. loop-audit reads it:

Result Meaning Effect on the score
proven A drill caught its fault and a benign case got through, and nothing failed Counts. The gate being proven is required for L3
failed A drill failed That guardrail's points are withdrawn: gate → gateYaml, verifier → verifier, breaker → stall detection. L3 is blocked
untested Everything was skipped, or only one direction passed Not proven. For the gate, loop-gate couldn't drill it, so gateYaml is withdrawn
stale Drilled against a different gate.yaml, or a canary over 30 days old Ignored until re-recorded

A few design points:

  • Both directions are required, as loop-drill already requires. A denylist of ** catches every seeded fault, so sensitivity alone can't prove a gate.
  • The gate proof is bound to the policy. The record holds a sha256 of gate.yaml with line endings normalised, so Windows and Linux checkouts agree. Weakening the policy after recording drops L3 until the drills run again. Gate and breaker drills are deterministic, so they don't otherwise expire. Verifier and injection canaries do.
  • A partial run doesn't lose proof. --only verifier --record replaces the verifier group and keeps the gate and breaker results.
  • Failures are recorded too. A verifier the canary caught approving seeded defects stops counting as a verifier. That is Verifier Theater, and the audit now says so by name.
  • No new dependencies. loop-audit only reads JSON, so it doesn't depend on loop-drill. That matters because loop-drill isn't on npm yet (see below).

3. The reference repo proves its own guardrails

  • loop-drill.json is committed. The reference repo stays at 100/L3, now with Guardrails proven by loop-drill: breaker, gate.
  • ci-audit-gates.sh fails if that record is missing, stale, failing, or doesn't prove the gate, and it prints the command to re-record. ci-validate-gates.sh already re-runs the drills against the live gate.yaml. Together they mean the committed record can't claim more than the drills show, and editing gate.yaml without re-recording fails CI.

Also

  • A starter fix. The changelog-drafter starter's .claude/agents/verifier.md had no frontmatter, so Claude Code would never load it as a subagent. I gave it the name and description from its Codex twin. The starter's score doesn't change, because the Codex verifier already counted.
  • Docs. The loop-audit README gets a Present is not proven section, and the loop-drill README gets Recording proof for loop-audit. The loop-audit CHANGELOG has an Unreleased entry. The L3 rows in docs/QUICKSTART.md, docs/operating-loops.md and SECURITY.md are updated.

Effect on scores

Before and after for the reference repo, all 17 starters and the fake repo:

Target Before After
The 142-byte fake repo 100 L3 72 L1
Reference repo, before loop-drill.json was recorded 100 L3 100 L2
Reference repo, with the committed record 100 L3 100 L3
All 17 starters — unchanged

No real file in the repo is empty, and no real skill directory lacks a loadable SKILL.md. I checked all 102 SKILL.md and agent files; the only one lacking frontmatter was the starter fixed above.

Verification

  • Tests: loop-audit has 50 (20 new), loop-drill has 69 (12 new). New coverage includes:
    • The fake repo byte-for-byte.
    • Hollow skills and agents.
    • Every proof verdict.
    • CRLF vs LF fingerprints.
    • --only merging.
    • --json --record keeping stdout parseable.
    • A full repo reaching L3 with a record, then losing it when gate.yaml is weakened.
    • One sha256 test vector pinned in both suites, so the two tools can't drift on the fingerprint. I checked it against sha256sum.
  • Every guard is load-bearing. I removed each of 21 guards in turn and confirmed a test fails every time. They include:
    • The empty-file check, {} JSON, frontmatter and the description requirement, the gate shape and .github content checks.
    • L3 needing proof, L3 blocked by any failure, and each of the verifier, gate-failed, gate-untested and breaker withdrawals.
    • Gate hash binding, the gate filename, canary expiry, CRLF normalisation on both sides, the both-directions rule, the schema check, merge keeping groups, and skip reasons being kept.
  • Both required workflows pass locally.
    • ci-validate-gates.sh exits 0 with 362 tests passing and none failing, including the loop-drill dogfood step.
    • ci-audit-gates.sh passes: Reference score: 100, Reference guardrails proven: breaker, gate.
    • before-after-demo.sh goes 7 → 64 → 94 (L2), as before.
  • Two things I had to work around on Windows, neither caused by this PR:

Limits — please read

  • Valid is not the same as good. A 394-byte repo of technically valid files still reaches 100/L3: one-line skills with frontmatter, a two-line gate.yaml that really blocks .env, keyword-stuffed LOOP.md, plus a real loop-drill --record. Its gate genuinely works, so the proof is true, but nothing judges whether the skills' instructions are any good. That limit is inherent to scoring files. The next lever would be requiring a recorded verifier canary for L3. I haven't done that, because it costs real verifier runs and would need a verifier command for this repo. That's your call.
  • The record is a claim the repo makes about itself, like Last run:. Someone determined can hand-write one with the right sha256. It raises the bar from touch to deliberate forgery, and CI re-running the drills (as this repo now does) is what keeps it honest.

Merge order: loop-drill must be on npm before loop-audit is released

@cobusgreyling/loop-drill isn't published yet: npm view returns 404, and there's no release-loop-drill.yml. Once this ships in a loop-audit release, L3 needs a record that users can only produce with npx @cobusgreyling/loop-drill. So I've not bumped loop-audit's version. The CHANGELOG entry is under Unreleased, so you can release it after loop-drill's first publish. Merging on its own is safe: it only changes this repo's code and CI.

Overlap with #641

#641 adds --agent-cmd and an injection drill to loop-drill's CLI, and this PR adds --record to the same file. Whichever merges second will have a small, mechanical conflict in cli.ts, and I'll rebase it. Grouping is by drill-id prefix, so injection results get recorded and scored with no further change. A failing injection canary then blocks L3 like any other failing guardrail.

…er it needs to load

Claude Code only loads a subagent file with name and description
frontmatter. This one had none, so the starter's Claude verifier was never
invocable. Name and description come from its Codex twin.
Results are grouped by guardrail (gate, breaker, verifier, ...). A run
replaces only the groups it drilled, so refreshing the cheap offline drills
doesn't discard a slow canary. Failures and skip reasons are recorded too.

The gate group carries a sha256 of the policy file, with line endings
normalised so Windows and Linux checkouts agree. loop-audit uses it to tell
a proof of the current gate.yaml from a proof of an older one.
A repo of 19 empty files (142 bytes) scored 100/L3, the same as the
reference repo: nearly every signal was fileExists(), and a skill counted if
its directory existed.

Placeholders no longer score. Empty or whitespace-only files, and {} / []
JSON, earn nothing and are listed under "Not counted". A skill or Claude
verifier agent needs name + description frontmatter to count. gate.yaml
needs version: 1 and a denylist. .github/ needs a non-empty file.

L3 now requires proven guardrails, read from loop-drill.json: the current
gate.yaml must pass its drills in both directions, and no recorded guardrail
may be failing. A guardrail loop-drill shows failing loses its points (gate,
verifier, breaker). Proof goes stale when gate.yaml changes, or after 30
days for canaries.

The fake repo now scores 72/L1. No starter's score changes.
Commit loop-drill.json so the reference repo stays at L3 on proof rather
than files. ci-audit-gates.sh now fails when the record is missing, stale
or failing, and prints the command to re-record. ci-validate-gates.sh
already re-runs the drills against the live gate.yaml, so the committed
record cannot claim more than they show.

The L3 criteria in QUICKSTART, operating-loops and SECURITY now mention
the proof.
@github-actions

Copy link
Copy Markdown
Contributor

Thanks @THRISHAL12345 for contributing a docs improvement — visible, reviewable PRs like this grow the reference for everyone.

What happens next

  • Maintainer aims for same-day review on story, adopter, and scoped docs/example PRs (CONTRIBUTING.md).
  • good first issue PRs: comment on the linked issue so we can assign and close on merge.

More ways to help

— loop-engineering maintainers

@github-actions

Copy link
Copy Markdown
Contributor

This PR changes paths that must run the real validate and audit workflows (tools, patterns, scripts, or CI).

Fork PRs from first-time contributors start with those workflows waiting for approval. A maintainer needs to open the Checks tab and click Approve and run workflows. Until that happens, branch protection will show the PR as blocked even after a review.

Content-only PRs (docs/, examples/, stories/, skills/, root markdown) skip this step — required checks are posted from this workflow instead.

— loop-engineering fork-pr-gate

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant