Vector compression with TurboQuant codecs for embeddings, retrieval, and KV-cache. 10x compression, pure NumPy core — optional GPU acceleration via PyTorch (CUDA/MPS) or MLX (Metal).
-
Updated
Sep 13, 2026 - Python
Vector compression with TurboQuant codecs for embeddings, retrieval, and KV-cache. 10x compression, pure NumPy core — optional GPU acceleration via PyTorch (CUDA/MPS) or MLX (Metal).
GenPark AI Agent Skill - Uniform 8-bit scalar quantization (SQ8) for high-density embedding compression and asymmetric dot product distance.
GenPark AI Agent Skill - Uniform 8-bit scalar quantization (SQ8) for high-density embedding compression and asymmetric dot product distance.
Product Quantization (PQ) vector compression engine partitioning embeddings into discrete codebook indices
Product Quantization (PQ) vector compression engine partitioning embeddings into discrete codebook indices
QJL sign-based vector compression and scoring in Rust — near-optimal distortion rate, append-only persistence, CPU-only, no LLM
Source code for ICML'26 paper "RQ-MoE: Residual Quantization via Mixture of Experts for Efficient Input-Dependent Vector Compression".
TensoRAG: High-performance Python engine for vector compression & fast RAG search via ML-GSVD. Reduces RAM up to 91%, speeds up search 14x. Independent, private hobby project implementing the PhD findings of Dr. L. Khamidullina (TU Ilmenau); no official affiliation with the author or university.
Accelerate LLM KV cache compression with a PyTorch TurboQuant implementation for efficient, high-quality vector quantization.
CommitMind: Semantic search for Git commit history powered by TurboQuant vector compression (ICLR 2026). Search commits by meaning, not just keywords.
AI-powered log anomaly detection CLI — learns normal patterns, detects anomalies with semantic embeddings, matches past incidents. Powered by TurboQuant 3-bit compression (ICLR 2026).
ChatMind: Semantic search for Discord & KakaoTalk chat messages. Search by meaning, not keywords. Powered by TurboQuant compression (ICLR 2026).
From-scratch survey of vector compression (JL, SVD, sign/scalar quantization, TurboQuant) evaluated on RAG retrieval over a linear-algebra textbook
To associate your repository with the vector-compression topic, visit your repo's landing page and select "manage topics."