Description
audit_kg_integrity answers "which documents does this graph object belong to?" from the object's graph source_id alone:
# lightrag/tools/kg_integrity_repair.py:50-52
def _split_sources(record: dict[str, Any] | None) -> list[str]:
raw = (record or {}).get("source_id") or ""
return [chunk_id for chunk_id in raw.split(GRAPH_FIELD_SEP) if chunk_id]
entity_chunks and relation_chunks appear nowhere in the 348-line file. But source_id is a capped view — apply_source_ids_limit trims it to MAX_SOURCE_IDS_PER_ENTITY / MAX_SOURCE_IDS_PER_RELATION (200 each by default, SOURCE_IDS_LIMIT_METHOD=KEEP), and the full list lives in the tracking rows. docs/design/PurgeRecoveryContract.md states the order explicitly:
Chunk tracking outranks graph source_id. Within a surviving entity or relation, the entity_chunks / relation_chunks row is the authoritative chunk list; the graph node's source_id is only a truncated view of it (apply_source_ids_limit) […] Genuinely missing attribution is repaired by audit_kg_integrity, never by the incremental path.
So the component the contract names as the authority is the one that reads only the non-authoritative view. Two wrong outputs follow.
Under-reporting (unconditional). A document whose chunks fall outside an object's cap window is absent from doc_entities / doc_relations, so its anchor gap is not reported and apply=True writes it an incomplete anchor set.
False certification (the severe one). If all of a document's chunks fall outside, the audit classifies it as anchorless_docs — which its own docstring calls "not merely unproven but proven empty" — and apply=True writes it empty full_entities / full_relations rows. That present-and-empty pair is the anchors proof under the purge contract, so the next adelete_by_doc_id returns 200, purges 0 entities and 0 relations, and deletes the document's chunks. The attribution carrier goes, the objects it attributed stay — the governing invariant inverted:
a purge must never delete something that CARRIES attribution — a chunk row or an anchor row that names objects — and leave those objects behind.
The contract also forbids this exact input choice for the sibling repair tool (L38): "It never seeds from graph source_id, whose reuse is precisely the provenance downgrade this section forbids."
Scope
The severe path needs a document with no anchor row — pre-anchor legacy data, ainsert_custom_kg, or a lost anchor write — whose chunks fall outside the KEEP window. Under the default KEEP method the window holds the oldest ids, and merge Phase 0 writes anchors for every modern document before any truncation, so this is a large-legacy-corpus condition rather than every deployment. That population is precisely this tool's target audience: the operator sent here by the 409 refusal. The under-reporting half has no such precondition.
Steps to reproduce
Two documents both contribute ALICE; the graph view keeps only the first one's chunk, as the cap would:
await rag.text_chunks.upsert({
C1: {"content": "a", "full_doc_id": D1, "chunk_order_index": 0},
C2: {"content": "b", "full_doc_id": D2, "chunk_order_index": 0},
})
await rag.chunk_entity_relation_graph.upsert_node(
"ALICE", {"entity_id": "ALICE", "entity_type": "PERSON",
"description": "x", "source_id": C1} # capped view
)
await rag.entity_chunks.upsert({"ALICE": {"chunk_ids": [C1, C2], "count": 2}}) # authoritative
# D1 gets anchors, D2 does not (legacy document)
await rag.full_entities.upsert({D1: {"entity_names": ["ALICE"], "count": 1}})
await rag.full_relations.upsert({D1: {"relation_pairs": [], "count": 0}})
report = await audit_kg_integrity(rag, apply=False)
Output on main @ 6702eea:
图节点 ALICE 的 source_id : chunk-of-doc-one <- capped
entity_chunks[ALICE] : [chunk-of-doc-one, chunk-of-doc-two] <- authoritative
missing_entity_anchors : {} <- under-reported
anchorless_docs : ['doc-two'] <- certified "proven empty"
doc-two owns ALICE and was certified as owning nothing.
Expected Behavior
The scan resolves each object's provenance from the authoritative chunk list, so a document that owns graph objects is never certified empty and its anchor gap is reported with the real names — the behavior PurgeRecoveryContract.md already states:
Absence is only ever concluded from the completed scan; a document that does own graph objects is repaired with its real names, never blanked.
LightRAG Config Used
Any backend combination — the tool touches only BaseKVStorage.get_by_ids and BaseGraphStorage.get_all_nodes / get_all_edges. Defaults MAX_SOURCE_IDS_PER_ENTITY=200, MAX_SOURCE_IDS_PER_RELATION=200, SOURCE_IDS_LIMIT_METHOD=KEEP.
Logs and screenshots
Nothing is logged. apply=True reports Wrote empty recovery anchors for N document(s) with no graph contributions, and the later purge reports success.
Additional Information
- LightRAG Version: main @ 6702eea
- Operating System: macOS 15 (arm64); platform-independent
- Python Version: 3.13.5
- Related Issues: the fail-closed purge work that introduced the certification
Description
audit_kg_integrityanswers "which documents does this graph object belong to?" from the object's graphsource_idalone:entity_chunksandrelation_chunksappear nowhere in the 348-line file. Butsource_idis a capped view —apply_source_ids_limittrims it toMAX_SOURCE_IDS_PER_ENTITY/MAX_SOURCE_IDS_PER_RELATION(200 each by default,SOURCE_IDS_LIMIT_METHOD=KEEP), and the full list lives in the tracking rows.docs/design/PurgeRecoveryContract.mdstates the order explicitly:So the component the contract names as the authority is the one that reads only the non-authoritative view. Two wrong outputs follow.
Under-reporting (unconditional). A document whose chunks fall outside an object's cap window is absent from
doc_entities/doc_relations, so its anchor gap is not reported andapply=Truewrites it an incomplete anchor set.False certification (the severe one). If all of a document's chunks fall outside, the audit classifies it as
anchorless_docs— which its own docstring calls "not merely unproven but proven empty" — andapply=Truewrites it emptyfull_entities/full_relationsrows. That present-and-empty pair is theanchorsproof under the purge contract, so the nextadelete_by_doc_idreturns 200, purges 0 entities and 0 relations, and deletes the document's chunks. The attribution carrier goes, the objects it attributed stay — the governing invariant inverted:The contract also forbids this exact input choice for the sibling repair tool (L38): "It never seeds from graph
source_id, whose reuse is precisely the provenance downgrade this section forbids."Scope
The severe path needs a document with no anchor row — pre-anchor legacy data,
ainsert_custom_kg, or a lost anchor write — whose chunks fall outside the KEEP window. Under the default KEEP method the window holds the oldest ids, and merge Phase 0 writes anchors for every modern document before any truncation, so this is a large-legacy-corpus condition rather than every deployment. That population is precisely this tool's target audience: the operator sent here by the 409 refusal. The under-reporting half has no such precondition.Steps to reproduce
Two documents both contribute
ALICE; the graph view keeps only the first one's chunk, as the cap would:Output on
main@ 6702eea:doc-twoownsALICEand was certified as owning nothing.Expected Behavior
The scan resolves each object's provenance from the authoritative chunk list, so a document that owns graph objects is never certified empty and its anchor gap is reported with the real names — the behavior
PurgeRecoveryContract.mdalready states:LightRAG Config Used
Any backend combination — the tool touches only
BaseKVStorage.get_by_idsandBaseGraphStorage.get_all_nodes/get_all_edges. DefaultsMAX_SOURCE_IDS_PER_ENTITY=200,MAX_SOURCE_IDS_PER_RELATION=200,SOURCE_IDS_LIMIT_METHOD=KEEP.Logs and screenshots
Nothing is logged.
apply=TruereportsWrote empty recovery anchors for N document(s) with no graph contributions, and the later purge reports success.Additional Information