Repository navigation
docs: accuracy sweep — misattributed DOIs, a guard that could not fail, and count drift - #323
Merged
Merged
Conversation
…l, count drift Three classes of defect, all live in a public repository. 1. Two DOIs in the README research table belonged to other researchers. 10.5281/zenodo.15105866 is a MALDI mass-spectrometry dataset by Ranes et al.; 10.5281/zenodo.15106553 is an e-learning article by Toshtemirov. Both were published here under Saleme titles, in four files each - including docs/proposals/attestation-schema-proposal.md, which is written for a standards venue. All nine DOIs in the repo were re-verified by content negotiation against doi.org; seven resolve to Saleme records and are kept. The two were replaced with verified records rather than re-pointed at a guessed identifier. The README carries a standing correction. 2. test_test_count_consistent_in_crosswalk could not fail. It matched "(\d+) security tests" against cli.py, a string cli.py does not contain, so the assertion was unreachable behind a truthiness check. It passed while the crosswalk said 595 and the canonical count was 603. Same defect class as the v4.13.1 HITL bug: a check that reports safety it never performed. 3. "595 tests / 43 modules" in twelve live files, including both AIUC-1 submission documents. README, SKILL.md and TEST-INVENTORY were guarded and correct; nothing guarded the rest. EVALUATION_PROTOCOL.md also claimed a 130-test suite and a "Framework version: 3.1 (189 tests)" footer. Adds repo-wide count guards over every live document, with dated snapshots excluded so history is not rewritten. The first version of the matcher missed the parenthetical "(603 tests)" form and passed a planted stale count; a plant test now pins that it fires on six real phrasings and stays silent on x402/L402/Ed25519. README and ROADMAP now state that a 2026-08-02 OpenAlex audit found 30 citation edges and 0 qualifying independent citations. ROADMAP had described the foundation as "peer-cross-citing DOIs", asserting external citation the audit disproves. Comparison table: Invariant Labs' mcp-scan redirects to snyk/agent-scan and was counted as two competitors. MCP coverage restated as 46; enterprise platforms corrected from 20 to 58. ROADMAP release history extended v4.4 -> v4.13.1 with known gaps stated. 361 tests pass (was 358). Count 603/44 unchanged. OWASP validator passes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes and found 1 potential issue.
❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, have a team admin enable autofix in the Cursor dashboard.
Reviewed by Cursor Bugbot for commit 7927afa. Configure here.
Bugbot on #323: the new sweep walked only .md/.py/.toml, so CITATION.cff was never checked. It still said "595 executable security tests across 43 modules" and version 4.10.0 while the canonical count was 603/44 and the release was 4.13.1 -- and that is the file GitHub reads for "Cite this repository", so CI could stay green while public citation metadata was wrong. A guard written to catch stale public claims had a blind spot at the most public claim in the repository. Extension list now includes .cff and .rst. Both guards verified against a planted stale count in CITATION.cff. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.

Reopened against
mainafter #321 merged (the original #322 auto-closed when its base branch was deleted). Same commit, rebased.Asked to make the repo accurate and current. Three classes of defect turned up, all live in a public repository.
1. Two DOIs belonged to other researchers
10.5281/zenodo.1510586610.5281/zenodo.15106553Verified twice — Zenodo record API and
doi.orgcontent negotiation.They appeared in four files each, including
docs/proposals/attestation-schema-proposal.md, which is written for submission to a standards venue. Citing another researcher's mass-spectrometry dataset as your own normalization-of-deviance paper, in a document arguing for a shared evidence schema, is the kind of thing that ends a conversation with a standards body.All nine DOIs in the repo were re-verified. Seven resolve to Saleme records and are retained. The two above were replaced with verified records rather than re-pointed at a guessed identifier — no Zenodo record under either title by this author was located, and inventing one would repeat the original error. The README carries a standing correction naming the real authors.
2. A guard that could not fail
test_test_count_consistent_in_crosswalkpassed while the crosswalk said 595 and the canonical count was 603. Two faults worth naming: it compared two documents to each other instead of to the source of truth, and it guarded the comparison behind a truthiness check, so the guard's own failure looked like a pass.Same defect class as the HITL bug in #321 — a check that reports safety it never performed. Different subsystem, same root.
3.
595 tests / 43 modulesin twelve live filesIncluding both AIUC-1 submission documents, the docs index, QUICKSTART, STRATEGY, the launch posts,
CLAUDE.mdandfree_scan.py. README, SKILL.md and TEST-INVENTORY were guarded and correct. Guarding three files did not guard the repository — and the ones that drifted are what outside readers see.Also found:
EVALUATION_PROTOCOL.mdclaimed a 130-test suite and carried aFramework version: 3.1 (189 tests)footer.The new guards, and why they are trustworthy
test_no_stale_test_count_anywhereandtest_no_stale_module_count_anywheresweep every live document. Dated snapshots — CHANGELOG, evaluation reports, archived roadmaps, blog posts — are excluded so history is not rewritten to today's number.The first version of the matcher was itself broken. It missed the parenthetical
(603 tests)form used in QUICKSTART, so a planted stale count sailed through. Caught by planting one rather than assuming.test_the_count_guard_can_actually_failnow pins that the matcher fires on six real phrasings and stays silent onx402 tests,L402 testsandEd25519.Verified by planting stale counts in three real files:
Also corrected
invariantlabs-ai/mcp-scannow redirects tosnyk/agent-scan; Invariant and Snyk were counted as two separate competitors. Merged, star counts re-verified (Snyk 2.9K, Cisco 1.0K, Garak 8.7K).v4.2; nowv4.13.1, with a note that unreachable and unserviced targets report INCONCLUSIVE.Verification
count_tests.py→ 603 across 44 modules, unchangedvalidate_owasp_agentic_mapping.py→ PASS🤖 Generated with Claude Code
Note
Low Risk
Docs and test-guard changes only; no production harness logic. Main risk is over-broad regex guards flagging benign phrases, mitigated by
test_the_count_guard_can_actually_fail.Overview
Documentation accuracy sweep fixes misattributed research citations, stale 603 tests / 44 modules claims across many live docs, and a crosswalk test that could never fail—without changing harness behavior.
Research citations: Two Zenodo DOIs that pointed at other researchers’ work are removed from README, AIUC-1 submission materials, blog posts, and the attestation schema proposal; they are replaced with verified Saleme records where applicable, plus a standing README correction and explicit OpenAlex citation honesty (0 independent citations).
Counts and metadata:
595/43(and outdated suite sizes inEVALUATION_PROTOCOL.md) are updated to 603/44 in AIUC-1 docs, QUICKSTART,CITATION.cff,CLAUDE.md,free_scan.py, launch posts, and related files.CITATION.cffis bumped to v4.13.1.README / ROADMAP: Competitor table dedupes Invariant → Snyk Agent Scan and refreshes star counts; MCP/enterprise coverage numbers are corrected; human oversight and OWASP v1.1 are added; release history runs through v4.13.1 with known gaps; positioning drops overstated “independent” / “peer-cross-citing” language.
Tests:
test_test_count_consistent_in_crosswalknow usesscripts/count_tests.py. New repo-wide guards scan live.md/.py/.toml/.cff/.rstfor stale totals (excluding dated snapshots), plustest_the_count_guard_can_actually_failso matchers cannot silently regress.Reviewed by Cursor Bugbot for commit d71e001. Bugbot is set up for automated code reviews on this repo. Configure here.