A fast, lightweight and easy-to-use Python library for splitting text into semantically meaningful chunks.
-
Updated
Jun 13, 2026 - Python
A fast, lightweight and easy-to-use Python library for splitting text into semantically meaningful chunks.
A package for parsing PDFs and analyzing their content using LLMs.
🍱 Semantically create chunks from large document for passing to LLM workflows
🧠✂️ SemanticSlicer — A smart text chunker for LLM-ready documents.
Embedding-driven, context-aware text chunking for Semantic Kernel and RAG workflows in .NET
⚡ Debug your RAG pipeline without leaving the terminal. Real-time chunking visualization, batch testing, quality metrics, and one-click export to LangChain/LlamaIndex.
This project is designed to extract text from documents and prepare it for processing by Large Language Models (LLM). Implemented a feature to store and utilize text style information, enabling the program to identify and segment content based on potential headers and titles.
Cutting-edge tool designed to intelligently segment text documents into optimally-sized chunks
Turn any document into a powerful Anki deck with NeuralDeck. This offline desktop app uses local AI to create high-quality flashcards from your PDFs, Word documents, and more. With smart deck matching, AI editing, and direct Anki sync, NeuralDeck is built for serious students who demand control, privacy, and efficiency.
A lightweight TypeScript text splitter for RAG applications
An exploration of text splitting and chunking in JavaScript
⚡ The fastest semantic text chunking library a SIMD-accelerated Rust core with a sync + async Python API.
A service-oriented .NET library for AI with interchangeable orchestrations and vector stores.
A RAG (Retrieval-Augmented Generation) pipeline implementation using .NET 10, Gemini AI, and Qdrant Vector Database.
Sementic chunking algorithm in (mostly) Go
Preprocess document service for RAG (Retriveal Augumented Generation)
Text Chunking Strategies for RAG — a practical Python demo showcasing fixed-size, sentence-based, and recursive chunking techniques for LLM embeddings, vector databases, and Retrieval-Augmented Generation (RAG) pipelines.
A Streamlit-based semantic search engine that converts documents into embeddings and retrieves the most meaningful text chunks using cosine similarity and dynamic chunking.
Smart text chunker for LLM preprocessing (sections → paragraphs → sentences → hard splits).
Chinese-aware RAG text chunking with exact source offsets, Markdown structure, and LangChain integration.
To associate your repository with the text-chunking topic, visit your repo's landing page and select "manage topics."