Skip to content
#

multi-gpu

Here are 121 public repositories matching this topic...

Self-hosted private AI workspace on owned Ubuntu hardware. llama.cpp layer-splits Qwen3.8-27B across an RTX 3060 and RTX 3070; a FastAPI app adds login-gated streaming chat, SearXNG search, page fetch, weather, Hugging Face and optional document retrieval. No LAN or public listener - remote access via Cloudflare Tunnel. AGPL-3.0.

  • Updated Sep 5, 2026
  • Python

Simulated Multi-GPU inference engine implementing Megatron-style Tensor Parallelism, GPipe Pipeline Parallelism, and KV Cache Sharding from scratch. Features bit-identical fidelity validation on Qwen2-0.5B weights and analytical communication cost modeling.

  • Updated Jul 29, 2026
  • Python

Quality-first local MiniMax H3 video-series generation and loopback API for dual RTX 4090 workstations, with native audio, references, P8/P9 continuity, and preserved artifacts.

  • Updated Sep 4, 2026
  • Python

Add this topic to your repo

To associate your repository with the multi-gpu topic, visit your repo's landing page and select "manage topics."

Learn more