Skip to content

[Feature Request]:Expose graph extraction and query APIs as stable public interface #3919

Description

@LXXiaogege

Do you need to file a feature request?

  • I have searched the existing feature request and this feature request is not already filed.
  • I believe this is a legitimate feature request, not just a question or bug.

Feature Request Description

Hi LightRAG team,

Thanks for the excellent work on this library! We're using LightRAG in a production RAG system and have found that the graph extraction and query capabilities are extremely valuable. However, our architecture requires more fine-grained control than the all-in-one LightRAG.insert() and LightRAG.query() methods provide.

Current Usage

We're currently calling internal functions directly to achieve our needs:

from lightrag.operate import extract_entities, merge_nodes_and_edges
from lightrag.kg import kg_query

During indexing

entities, relationships = await extract_entities(chunks, ...)
merged = await merge_nodes_and_edges(entities, relationships, ...)

During retrieval

results = await kg_query(query, mode="hybrid", ...)

This works well, but these functions are not documented as public APIs, which creates uncertainty around version upgrades and long-term maintenance.

Use Case

Our system has the following requirements that prevent us from using the built-in insert() / query() workflow:

  1. Custom chunking strategy: We use domain-specific chunkers with parent-child relationships, semantic boundaries, and metadata enrichment that happens before graph extraction.
  2. Separate vector store: We already have a Qdrant-based dense+sparse hybrid retrieval pipeline. We only need LightRAG's graph extraction and graph query capabilities, not its built-in chunk embedding or vector storage.
  3. Multi-stage retrieval fusion: Graph results need to merge with vector search, full-text search, and reranking in a unified pipeline, rather than having LightRAG generate the final answer.
  4. Batch control and progress tracking: We need fine-grained control over batch sizes, progress callbacks, and error recovery during large-scale indexing (10K+ documents).

Proposal

Would you consider exposing a stable, documented API for graph-only operations? Something like:

from lightrag import GraphExtractor, GraphStore, GraphQuery

Extraction

extractor = GraphExtractor(llm=..., global_config=...)
entities, relationships = await extractor.extract(chunks)

Storage (bring your own vector store for entity/relationship embeddings)

graph_store = GraphStore(...)
await graph_store.merge(entities, relationships)

Query

query_engine = GraphQuery(graph_store=..., llm=...)
results = await query_engine.query(question, mode="hybrid", top_k=20)

This would allow advanced users to:

  • Integrate LightRAG's graph capabilities into existing RAG pipelines
  • Use custom embedding models and vector stores while leveraging your graph extraction logic
  • Control indexing/query workflows at a granular level
  • Upgrade LightRAG with confidence that the graph-only API remains stable

Benefits

  • Broader adoption: Many production RAG systems already have retrieval infrastructure but need better graph capabilities.
  • Clearer separation of concerns: Graph extraction, storage, and query as distinct stages make the library easier to understand and extend.
  • Backward compatibility: The current all-in-one API can remain unchanged as a convenience wrapper.

Current Workaround

For now, we've pinned to lightrag-hku==1.5.7 and wrapped the internal functions in our own adapter layer with comprehensive integration tests. This works, but requires careful validation before any version upgrade.

Let me know if this aligns with the project's roadmap! Happy to provide more details about our use case or contribute to API design discussions.

Thanks again for the great work!

Additional Context

No response

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions