low-bit
Here are 16 public repositories matching this topic...
Phonon: open speech recognition models (Phonon-2, Phonon-1) — CLI, CPU and CUDA images
-
Updated
Oct 1, 2026 - Python
Get more intelligence from every bit. Better quantization formats and smarter calibration let larger, stronger models run smoothly on the hardware you already own.
-
Updated
Oct 2, 2026 - C++
QuantFace: Towards Lightweight Face Recognition by Synthetic Data Low-bit Quantization
-
Updated
Jun 26, 2022 - Python
admm for cnn layerwise weight low bit quantization
-
Updated
Sep 25, 2019 - Python
Memory-efficient PyTorch optimizers with Adaptive Log-Space quantization, fused CUDA kernels, and Adafactor, CAME, APOLLO support.
-
Updated
Sep 9, 2026 - Python
2.898-BPW Qwen3-8B with direct-packed CPU/CUDA inference
-
Updated
Aug 7, 2026 - Python
Efficient low-bit KV-cache compression research with honest metadata accounting.
-
Updated
Apr 24, 2026 - Python
Pure-Julia CPU inference engine for BitNet b1.58 ternary LLMs
-
Updated
Jun 17, 2026 - Julia
Microscaling (MX) low-bit quantization formats, implemented and benchmarked on VLMs.
-
Updated
Aug 16, 2026 - Jupyter Notebook
Compute-for-memory research lab for running larger LLMs on memory-constrained consumer GPUs.
-
Updated
Sep 23, 2026
Quantization toolkit for large language models: run and store huge models on small VRAM. Inspired by Unsloth's dynamic quantization.
-
Updated
Sep 11, 2026
qpeft: quantization-aware training (QAT) and PEFT for LLMs in PyTorch. EfficientQAT, QA-LoRA and PEQA for Hugging Face models at 2/3/4 bits: train scales, zero-points or LoRA adapters, then merge into a low-bit integer model (GPTQ layout) that is tested to equal the trained one. Pure-torch and torchao backends.
-
Updated
Sep 29, 2026 - Python
-
Updated
Aug 16, 2026 - Jupyter Notebook
Add this topic to your repo
To associate your repository with the low-bit topic, visit your repo's landing page and select "manage topics."