Skip to content

[Bug]: audit_kg_integrity resolves provenance from the capped source_id, not chunk tracking #4083

Description

@JessYanCoding

Description

audit_kg_integrity answers "which documents does this graph object belong to?" from the object's graph source_id alone:

# lightrag/tools/kg_integrity_repair.py:50-52
def _split_sources(record: dict[str, Any] | None) -> list[str]:
    raw = (record or {}).get("source_id") or ""
    return [chunk_id for chunk_id in raw.split(GRAPH_FIELD_SEP) if chunk_id]

entity_chunks and relation_chunks appear nowhere in the 348-line file. But source_id is a capped view — apply_source_ids_limit trims it to MAX_SOURCE_IDS_PER_ENTITY / MAX_SOURCE_IDS_PER_RELATION (200 each by default, SOURCE_IDS_LIMIT_METHOD=KEEP), and the full list lives in the tracking rows. docs/design/PurgeRecoveryContract.md states the order explicitly:

Chunk tracking outranks graph source_id. Within a surviving entity or relation, the entity_chunks / relation_chunks row is the authoritative chunk list; the graph node's source_id is only a truncated view of it (apply_source_ids_limit) […] Genuinely missing attribution is repaired by audit_kg_integrity, never by the incremental path.

So the component the contract names as the authority is the one that reads only the non-authoritative view. Two wrong outputs follow.

Under-reporting (unconditional). A document whose chunks fall outside an object's cap window is absent from doc_entities / doc_relations, so its anchor gap is not reported and apply=True writes it an incomplete anchor set.

False certification (the severe one). If all of a document's chunks fall outside, the audit classifies it as anchorless_docs — which its own docstring calls "not merely unproven but proven empty" — and apply=True writes it empty full_entities / full_relations rows. That present-and-empty pair is the anchors proof under the purge contract, so the next adelete_by_doc_id returns 200, purges 0 entities and 0 relations, and deletes the document's chunks. The attribution carrier goes, the objects it attributed stay — the governing invariant inverted:

a purge must never delete something that CARRIES attribution — a chunk row or an anchor row that names objects — and leave those objects behind.

The contract also forbids this exact input choice for the sibling repair tool (L38): "It never seeds from graph source_id, whose reuse is precisely the provenance downgrade this section forbids."

Scope

The severe path needs a document with no anchor row — pre-anchor legacy data, ainsert_custom_kg, or a lost anchor write — whose chunks fall outside the KEEP window. Under the default KEEP method the window holds the oldest ids, and merge Phase 0 writes anchors for every modern document before any truncation, so this is a large-legacy-corpus condition rather than every deployment. That population is precisely this tool's target audience: the operator sent here by the 409 refusal. The under-reporting half has no such precondition.

Steps to reproduce

Two documents both contribute ALICE; the graph view keeps only the first one's chunk, as the cap would:

await rag.text_chunks.upsert({
    C1: {"content": "a", "full_doc_id": D1, "chunk_order_index": 0},
    C2: {"content": "b", "full_doc_id": D2, "chunk_order_index": 0},
})
await rag.chunk_entity_relation_graph.upsert_node(
    "ALICE", {"entity_id": "ALICE", "entity_type": "PERSON",
              "description": "x", "source_id": C1}      # capped view
)
await rag.entity_chunks.upsert({"ALICE": {"chunk_ids": [C1, C2], "count": 2}})   # authoritative
# D1 gets anchors, D2 does not (legacy document)
await rag.full_entities.upsert({D1: {"entity_names": ["ALICE"], "count": 1}})
await rag.full_relations.upsert({D1: {"relation_pairs": [], "count": 0}})

report = await audit_kg_integrity(rag, apply=False)

Output on main @ 6702eea:

图节点 ALICE 的 source_id     : chunk-of-doc-one          <- capped
entity_chunks[ALICE]          : [chunk-of-doc-one, chunk-of-doc-two]   <- authoritative
missing_entity_anchors        : {}                        <- under-reported
anchorless_docs               : ['doc-two']               <- certified "proven empty"

doc-two owns ALICE and was certified as owning nothing.

Expected Behavior

The scan resolves each object's provenance from the authoritative chunk list, so a document that owns graph objects is never certified empty and its anchor gap is reported with the real names — the behavior PurgeRecoveryContract.md already states:

Absence is only ever concluded from the completed scan; a document that does own graph objects is repaired with its real names, never blanked.

LightRAG Config Used

Any backend combination — the tool touches only BaseKVStorage.get_by_ids and BaseGraphStorage.get_all_nodes / get_all_edges. Defaults MAX_SOURCE_IDS_PER_ENTITY=200, MAX_SOURCE_IDS_PER_RELATION=200, SOURCE_IDS_LIMIT_METHOD=KEEP.

Logs and screenshots

Nothing is logged. apply=True reports Wrote empty recovery anchors for N document(s) with no graph contributions, and the later purge reports success.

Additional Information

  • LightRAG Version: main @ 6702eea
  • Operating System: macOS 15 (arm64); platform-independent
  • Python Version: 3.13.5
  • Related Issues: the fail-closed purge work that introduced the certification

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions