Skip to content

feat(tools): add bounded count-aware VDB consistency checks #4058

Description

@danielaskdd

Motivation

lightrag-rebuild-vdb already performs a forward source-to-vector existence check, but the check materializes all graph nodes and edges in client memory. We do not need a full bidirectional vector-ID inventory to decide whether an operator should rebuild a derived VDB.

For each target, a bounded forward probe plus exact source/target counts is sufficient for the operational consistency decision:

  1. If any authoritative source record has no vector counterpart, the target is inconsistent.
  2. If every source record has a vector counterpart but the exact counts differ, the target is inconsistent.
  3. If the forward probe is complete and the exact counts agree, report the target consistent under LightRAG's existing ID assumptions.
  4. Any inconsistent target is repaired through the existing drop-and-rebuild operation.

This replaces the broader full-population census proposed in #3997 and avoids whole-vector inventory, reverse-orphan enumeration, and per-document attribution.

Scope

  • Rewrite entity and relationship checks in check_vdb_consistency() to use BaseGraphStorage.iter_labels() and iter_edges().
  • Keep client-side graph memory bounded by batch_size plus bounded report examples; do not collect a whole-population seen set.
  • Add an exact, namespace/workspace-scoped vector record count capability with an exact-or-raise contract. Transport errors, unavailable containers, and partial reads must never become zero.
  • Extend the check to text_chunks -> chunks_vdb using bounded KV-key iteration.
  • Return an explicit per-target status such as consistent, inconsistent, inconclusive, incompatible, or not_applicable.
  • Preserve the current embedding-space mismatch refusal and the offline/no-writers operating requirement.
  • Direct inconsistent results to the existing entity/relation/chunk rebuild targets; do not add a new repair writer.
  • Make entity and relationship rebuild source scans bounded where they currently use get_all_nodes() / get_all_edges() and whole-population payload dictionaries.

Accepted limitation: relationship IDs

Relationship vector IDs currently hash delimiter-free endpoint concatenation and accept a historical reverse-order candidate. The mapping is therefore not mathematically injective: distinct endpoint pairs can theoretically share a candidate ID. This is a legacy identity limitation with a very low practical incidence and is accepted for this check. Document the assumption and add a focused test so the limitation remains explicit; do not reintroduce a whole-population ID set solely to rule it out.

Required semantics

  • inconsistent is a positive finding and recommends rebuild.
  • consistent means the bounded forward probe completed without misses and exact source/target counts agree, subject to the documented existing ID assumptions.
  • A failed/unsupported count or incomplete point read is inconclusive/unavailable, never consistent and never a zero count.
  • Report examples remain capped even when the number of missing records is large.

Out of scope

  • Full VDB ID enumeration.
  • Listing reverse-only vector IDs.
  • Per-document attribution of vector gaps.
  • Integration into kg_integrity_repair.
  • Automatic repair or startup checks.
  • Redesigning or migrating the relationship vector-ID format.

Acceptance criteria

  • Memory used by graph/KV source enumeration is bounded independently of source population size.
  • Forward misses and exact count mismatches are detected for entities, relationships, and chunks.
  • Equal exact counts plus a complete forward probe report consistency under the documented ID assumptions.
  • Count/read failures fail loud and never produce a false consistent result.
  • Existing rebuild targets remain the only repair mechanism.
  • Regression tests cover empty stores, a missing vector, a reverse-only extra detected by count mismatch, incompatible embedding space, failed counts, incomplete reads, bounded iteration, legacy reverse relation IDs, and the accepted ambiguous relationship-ID example.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or requesttrackedIssue is tracked by project

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions