Skip to content

docs: accuracy sweep — misattributed DOIs, a guard that could not fail, and count drift - #323

Merged
msaleme merged 2 commits into
mainfrom
docs/accuracy-sweep
Aug 3, 2026
Merged

msaleme merged 2 commits into
mainfrom
docs/accuracy-sweep

Conversation

@msaleme

@msaleme msaleme commented Aug 3, 2026 •

Copy link
Copy Markdown
Owner

Reopened against main after #321 merged (the original #322 auto-closed when its base branch was deleted). Same commit, rebased.

Asked to make the repo accurate and current. Three classes of defect turned up, all live in a public repository.

1. Two DOIs belonged to other researchers

DOI Published here as What it actually is
10.5281/zenodo.15105866 Normalization of Deviance in Autonomous Agent Systems MALDI mass-spectrometry dataset, Ranes / Moore / Patterson / Nicholson, 2025-03-29
10.5281/zenodo.15106553 Cognitive Style Governance for Multi-Agent Deployments E-learning article in Uzbek, Toshtemirov, 2025-03-29

Verified twice — Zenodo record API and doi.org content negotiation.

They appeared in four files each, including docs/proposals/attestation-schema-proposal.md, which is written for submission to a standards venue. Citing another researcher's mass-spectrometry dataset as your own normalization-of-deviance paper, in a document arguing for a shared evidence schema, is the kind of thing that ends a conversation with a standards body.

All nine DOIs in the repo were re-verified. Seven resolve to Saleme records and are retained. The two above were replaced with verified records rather than re-pointed at a guessed identifier — no Zenodo record under either title by this author was located, and inventing one would repeat the original error. The README carries a standing correction naming the real authors.

2. A guard that could not fail

cli_counts = re.findall(r'(\d+)\s+security tests', f.read())   # cli.py has no such string
if counts_in_crosswalk and cli_counts:                          # -> always False
    ...                                                         # -> never ran

test_test_count_consistent_in_crosswalk passed while the crosswalk said 595 and the canonical count was 603. Two faults worth naming: it compared two documents to each other instead of to the source of truth, and it guarded the comparison behind a truthiness check, so the guard's own failure looked like a pass.

Same defect class as the HITL bug in #321 — a check that reports safety it never performed. Different subsystem, same root.

3. 595 tests / 43 modules in twelve live files

Including both AIUC-1 submission documents, the docs index, QUICKSTART, STRATEGY, the launch posts, CLAUDE.md and free_scan.py. README, SKILL.md and TEST-INVENTORY were guarded and correct. Guarding three files did not guard the repository — and the ones that drifted are what outside readers see.

Also found: EVALUATION_PROTOCOL.md claimed a 130-test suite and carried a Framework version: 3.1 (189 tests) footer.

The new guards, and why they are trustworthy

test_no_stale_test_count_anywhere and test_no_stale_module_count_anywhere sweep every live document. Dated snapshots — CHANGELOG, evaluation reports, archived roadmaps, blog posts — are excluded so history is not rewritten to today's number.

The first version of the matcher was itself broken. It missed the parenthetical (603 tests) form used in QUICKSTART, so a planted stale count sailed through. Caught by planting one rather than assuming. test_the_count_guard_can_actually_fail now pins that the matcher fires on six real phrasings and stays silent on x402 tests, L402 tests and Ed25519.

Verified by planting stale counts in three real files:

planted '595 tests'       in docs/QUICKSTART.md  -> 1 failed
planted 'tests-595-'      in README.md           -> 1 failed
planted '43 test-bearing' in docs/README.md      -> 1 failed
restored                                         -> 2 passed

Also corrected

  • The comparison table listed the same repo twice. invariantlabs-ai/mcp-scan now redirects to snyk/agent-scan; Invariant and Snyk were counted as two separate competitors. Merged, star counts re-verified (Snyk 2.9K, Cisco 1.0K, Garak 8.7K).
  • MCP coverage restated as 46 (protocol 32 + supply-chain 4 + tool-poisoning repro 10); it said 31. Enterprise platforms corrected from "20" to 58 (core 31 + extended 27).
  • Citation honesty. README and ROADMAP now state that a 2026-08-02 OpenAlex audit found 30 citation edges and 0 qualifying independent citations — every edge a self-citation. ROADMAP had called the foundation "peer-cross-citing DOIs", asserting exactly the external citation the audit disproves, and claimed "independent, reproducible adversarial evidence" where adjudication is author-performed.
  • ROADMAP release history ran v3.9 → v4.4 with "Next" in progress, seven minor versions behind. Now v3.9 → v4.13.1 with dates reconciled against CHANGELOG, plus a stated known-gaps section (T16-S1, T10-S1/S3, one guidance-only control, author-performed adjudication).
  • The research-frontier list named intent-contract, multi-agent and memory work that shipped as modules. Replaced with the open research question for each, and the honest frontier: human-oversight measurement needs a study design, not more tests.
  • README: added a Human Oversight layer, an OWASP v1.1 row, and an Independent Reproduction entry crediting @VrtxOmega — scoped as one external party reproducing one pinned artifact, explicitly not endorsement or adoption.
  • Sample CLI output said v4.2; now v4.13.1, with a note that unreachable and unserviced targets report INCONCLUSIVE.

Verification

  • 361 tests pass (was 358)
  • count_tests.py → 603 across 44 modules, unchanged
  • validate_owasp_agentic_mapping.py → PASS
  • Every README link checked to resolve
  • No product behaviour changes; documentation, one dead test repaired, three added

🤖 Generated with Claude Code


Note

Low Risk
Docs and test-guard changes only; no production harness logic. Main risk is over-broad regex guards flagging benign phrases, mitigated by test_the_count_guard_can_actually_fail.

Overview
Documentation accuracy sweep fixes misattributed research citations, stale 603 tests / 44 modules claims across many live docs, and a crosswalk test that could never fail—without changing harness behavior.

Research citations: Two Zenodo DOIs that pointed at other researchers’ work are removed from README, AIUC-1 submission materials, blog posts, and the attestation schema proposal; they are replaced with verified Saleme records where applicable, plus a standing README correction and explicit OpenAlex citation honesty (0 independent citations).

Counts and metadata: 595/43 (and outdated suite sizes in EVALUATION_PROTOCOL.md) are updated to 603/44 in AIUC-1 docs, QUICKSTART, CITATION.cff, CLAUDE.md, free_scan.py, launch posts, and related files. CITATION.cff is bumped to v4.13.1.

README / ROADMAP: Competitor table dedupes Invariant → Snyk Agent Scan and refreshes star counts; MCP/enterprise coverage numbers are corrected; human oversight and OWASP v1.1 are added; release history runs through v4.13.1 with known gaps; positioning drops overstated “independent” / “peer-cross-citing” language.

Tests: test_test_count_consistent_in_crosswalk now uses scripts/count_tests.py. New repo-wide guards scan live .md/.py/.toml/.cff/.rst for stale totals (excluding dated snapshots), plus test_the_count_guard_can_actually_fail so matchers cannot silently regress.

Reviewed by Cursor Bugbot for commit d71e001. Bugbot is set up for automated code reviews on this repo. Configure here.

…l, count drift

Three classes of defect, all live in a public repository.

1. Two DOIs in the README research table belonged to other researchers.
   10.5281/zenodo.15105866 is a MALDI mass-spectrometry dataset by Ranes et
   al.; 10.5281/zenodo.15106553 is an e-learning article by Toshtemirov. Both
   were published here under Saleme titles, in four files each - including
   docs/proposals/attestation-schema-proposal.md, which is written for a
   standards venue. All nine DOIs in the repo were re-verified by content
   negotiation against doi.org; seven resolve to Saleme records and are kept.
   The two were replaced with verified records rather than re-pointed at a
   guessed identifier. The README carries a standing correction.

2. test_test_count_consistent_in_crosswalk could not fail. It matched
   "(\d+) security tests" against cli.py, a string cli.py does not contain,
   so the assertion was unreachable behind a truthiness check. It passed
   while the crosswalk said 595 and the canonical count was 603. Same defect
   class as the v4.13.1 HITL bug: a check that reports safety it never
   performed.

3. "595 tests / 43 modules" in twelve live files, including both AIUC-1
   submission documents. README, SKILL.md and TEST-INVENTORY were guarded and
   correct; nothing guarded the rest. EVALUATION_PROTOCOL.md also claimed a
   130-test suite and a "Framework version: 3.1 (189 tests)" footer.

Adds repo-wide count guards over every live document, with dated snapshots
excluded so history is not rewritten. The first version of the matcher missed
the parenthetical "(603 tests)" form and passed a planted stale count; a
plant test now pins that it fires on six real phrasings and stays silent on
x402/L402/Ed25519.

README and ROADMAP now state that a 2026-08-02 OpenAlex audit found 30
citation edges and 0 qualifying independent citations. ROADMAP had described
the foundation as "peer-cross-citing DOIs", asserting external citation the
audit disproves.

Comparison table: Invariant Labs' mcp-scan redirects to snyk/agent-scan and
was counted as two competitors. MCP coverage restated as 46; enterprise
platforms corrected from 20 to 58. ROADMAP release history extended v4.4 ->
v4.13.1 with known gaps stated.

361 tests pass (was 358). Count 603/44 unchanged. OWASP validator passes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, have a team admin enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 7927afa. Configure here.

Comment thread testing/test_code_quality.py
Bugbot on #323: the new sweep walked only .md/.py/.toml, so CITATION.cff was
never checked. It still said "595 executable security tests across 43 modules"
and version 4.10.0 while the canonical count was 603/44 and the release was
4.13.1 -- and that is the file GitHub reads for "Cite this repository", so CI
could stay green while public citation metadata was wrong.

A guard written to catch stale public claims had a blind spot at the most
public claim in the repository. Extension list now includes .cff and .rst.
Both guards verified against a planted stale count in CITATION.cff.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@msaleme
msaleme merged commit 19dbfb8 into main Aug 3, 2026
8 checks passed
@msaleme
msaleme deleted the docs/accuracy-sweep branch August 3, 2026 00:57
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant