Skip to content
#

quantization

Here are 64 public repositories matching this topic...

RTX 5090 Laptop (24GB) + ComfyUI: 12 model lines / 21 measured configs — MiniMax H3, Wan 2.2, Qwen-Image 2512, HunyuanVideo 1.5, FLUX.2 Klein, Z-Image-Turbo, LTX-Video, ACE-Step 1.5, YuE2, Stable Audio 3, Hunyuan3D 2.1, ERNIE-Image. A controlled A/B proves NVFP4 beats INT8. Bilingual 中文/English, licence-verified.

  • Updated Sep 22, 2026
  • HTML

A self-contained AI project that runs a quantized Large Language Model (Qwen2.5-0.5B) entirely on your local machine. Built with FastAPI and llama-cpp-python, this agent intelligently switches between standard chat and "Search Mode" to fetch real-time data from the internet. The project features a responsive HTML/CSS/JS frontend and is fully Docker

  • Updated Jan 8, 2026
  • HTML

An AI-powered MLOps assistant for effortless model compression. Upload PyTorch models to chat with a local LLM expert, receive hardware-aware optimization advice, and perform one-click FP16/INT8 quantization to reduce model size and latency.

  • Updated Sep 11, 2025
  • HTML

A bilingual field guide to Inference Engineering: the path from product constraints and model mechanics through hardware, runtimes, optimization techniques, modalities, and production operation.

  • Updated Aug 7, 2026
  • HTML

Add this topic to your repo

To associate your repository with the quantization topic, visit your repo's landing page and select "manage topics."

Learn more