Motivation
lightrag-rebuild-vdb already performs a forward source-to-vector existence check, but the check materializes all graph nodes and edges in client memory. We do not need a full bidirectional vector-ID inventory to decide whether an operator should rebuild a derived VDB.
For each target, a bounded forward probe plus exact source/target counts is sufficient for the operational consistency decision:
- If any authoritative source record has no vector counterpart, the target is inconsistent.
- If every source record has a vector counterpart but the exact counts differ, the target is inconsistent.
- If the forward probe is complete and the exact counts agree, report the target consistent under LightRAG's existing ID assumptions.
- Any inconsistent target is repaired through the existing drop-and-rebuild operation.
This replaces the broader full-population census proposed in #3997 and avoids whole-vector inventory, reverse-orphan enumeration, and per-document attribution.
Scope
- Rewrite entity and relationship checks in
check_vdb_consistency() to use BaseGraphStorage.iter_labels() and iter_edges().
- Keep client-side graph memory bounded by
batch_size plus bounded report examples; do not collect a whole-population seen set.
- Add an exact, namespace/workspace-scoped vector record count capability with an exact-or-raise contract. Transport errors, unavailable containers, and partial reads must never become zero.
- Extend the check to
text_chunks -> chunks_vdb using bounded KV-key iteration.
- Return an explicit per-target status such as
consistent, inconsistent, inconclusive, incompatible, or not_applicable.
- Preserve the current embedding-space mismatch refusal and the offline/no-writers operating requirement.
- Direct inconsistent results to the existing entity/relation/chunk rebuild targets; do not add a new repair writer.
- Make entity and relationship rebuild source scans bounded where they currently use
get_all_nodes() / get_all_edges() and whole-population payload dictionaries.
Accepted limitation: relationship IDs
Relationship vector IDs currently hash delimiter-free endpoint concatenation and accept a historical reverse-order candidate. The mapping is therefore not mathematically injective: distinct endpoint pairs can theoretically share a candidate ID. This is a legacy identity limitation with a very low practical incidence and is accepted for this check. Document the assumption and add a focused test so the limitation remains explicit; do not reintroduce a whole-population ID set solely to rule it out.
Required semantics
inconsistent is a positive finding and recommends rebuild.
consistent means the bounded forward probe completed without misses and exact source/target counts agree, subject to the documented existing ID assumptions.
- A failed/unsupported count or incomplete point read is
inconclusive/unavailable, never consistent and never a zero count.
- Report examples remain capped even when the number of missing records is large.
Out of scope
- Full VDB ID enumeration.
- Listing reverse-only vector IDs.
- Per-document attribution of vector gaps.
- Integration into
kg_integrity_repair.
- Automatic repair or startup checks.
- Redesigning or migrating the relationship vector-ID format.
Acceptance criteria
- Memory used by graph/KV source enumeration is bounded independently of source population size.
- Forward misses and exact count mismatches are detected for entities, relationships, and chunks.
- Equal exact counts plus a complete forward probe report consistency under the documented ID assumptions.
- Count/read failures fail loud and never produce a false consistent result.
- Existing rebuild targets remain the only repair mechanism.
- Regression tests cover empty stores, a missing vector, a reverse-only extra detected by count mismatch, incompatible embedding space, failed counts, incomplete reads, bounded iteration, legacy reverse relation IDs, and the accepted ambiguous relationship-ID example.
Motivation
lightrag-rebuild-vdbalready performs a forward source-to-vector existence check, but the check materializes all graph nodes and edges in client memory. We do not need a full bidirectional vector-ID inventory to decide whether an operator should rebuild a derived VDB.For each target, a bounded forward probe plus exact source/target counts is sufficient for the operational consistency decision:
This replaces the broader full-population census proposed in #3997 and avoids whole-vector inventory, reverse-orphan enumeration, and per-document attribution.
Scope
check_vdb_consistency()to useBaseGraphStorage.iter_labels()anditer_edges().batch_sizeplus bounded report examples; do not collect a whole-populationseenset.text_chunks -> chunks_vdbusing bounded KV-key iteration.consistent,inconsistent,inconclusive,incompatible, ornot_applicable.get_all_nodes()/get_all_edges()and whole-population payload dictionaries.Accepted limitation: relationship IDs
Relationship vector IDs currently hash delimiter-free endpoint concatenation and accept a historical reverse-order candidate. The mapping is therefore not mathematically injective: distinct endpoint pairs can theoretically share a candidate ID. This is a legacy identity limitation with a very low practical incidence and is accepted for this check. Document the assumption and add a focused test so the limitation remains explicit; do not reintroduce a whole-population ID set solely to rule it out.
Required semantics
inconsistentis a positive finding and recommends rebuild.consistentmeans the bounded forward probe completed without misses and exact source/target counts agree, subject to the documented existing ID assumptions.inconclusive/unavailable, never consistent and never a zero count.Out of scope
kg_integrity_repair.Acceptance criteria