multi-gpu
Here are 122 public repositories matching this topic...
Genuine multi-GPU FSDP full fine-tune of a 7B across 4x A100 (closes the scale gap), with the tied-embeddings, collective-save, and checkpoint-consolidation gotchas made concrete.
-
Updated
Jun 19, 2026 - Python
Self-hosted private AI workspace on owned Ubuntu hardware. llama.cpp layer-splits Qwen3.8-27B across an RTX 3060 and RTX 3070; a FastAPI app adds login-gated streaming chat, SearXNG search, page fetch, weather, Hugging Face and optional document retrieval. No LAN or public listener - remote access via Cloudflare Tunnel. AGPL-3.0.
-
Updated
Sep 5, 2026 - Python
Simulated Multi-GPU inference engine implementing Megatron-style Tensor Parallelism, GPipe Pipeline Parallelism, and KV Cache Sharding from scratch. Features bit-identical fidelity validation on Qwen2-0.5B weights and analytical communication cost modeling.
-
Updated
Jul 29, 2026 - Python
performance test of MNIST hand writings usign MXNet + TF
-
Updated
Jan 31, 2020 - Python
Lightweight terminal launcher and auto-optimizer for llm models using llama.cpp with hardware detection, tensor sharding, benchmarks, presets, context tuning, and OpenAI-compatible serving.
-
Updated
Sep 3, 2026 - Python
Experiments from 'NLP with Transformers' (fine-tuning, summarization, knowledge distillation) as standalone scripts, plus multi-GPU training with Hugging Face Accelerate: DDP gives ~1.8x the throughput of DataParallel on 2x RTX 4060 Ti
-
Updated
Sep 24, 2026 - Python
Layer split vs tensor parallelism in llama.cpp - measured on two platforms 16x apart in interconnect bandwidth
-
Updated
Oct 6, 2026 - Python
Reproducible PyTorch DDP scaling benchmark for NVIDIA GPUs. Measures weak/strong scaling, speedup, and parallel efficiency across single-GPU, multi-GPU, and multi-node configurations. Slurm-ready.
-
Updated
Jul 7, 2026 - Python
Multi-GPU video upscaling with frame-level parallelism
-
Updated
Jan 29, 2026 - Python
Linux-first dual RTX 3090 + Threadripper PRO AI server — Part 1 of a three-part series. Architecture, cost, airflow, build notes, and hardware documentation.
-
Updated
Aug 22, 2026 - Python
Run NVIDIA Tesla P40 (Pascal) + RTX 5070 Ti (Blackwell) together for llama.cpp — mixed NVIDIA drivers via KVM + AF_VSOCK RPC
-
Updated
Oct 6, 2026 - Python
Production-ready framework for training robust computer vision models. Features multi-GPU support, EMA tracking, label smoothing, and comprehensive robustness evaluation across 4 noise types. Includes scalable TF.Data pipeline, automated testing, Docker support, and CLI tools. Install: pip install robust-vision
-
Updated
Mar 13, 2026 - Python
Systematic VLA training optimization on 2× RTX 3090. WebDataset + FlashAttention-2 + FSDP → 3.3× throughput, 26% VRAM reduction. Profiler traces and W&B report linked. Reproducible in one command.
-
Updated
Jul 1, 2026 - Python
The generation engine behind Inline Studio: a typed graph backend that runs diffusion models on GPU, low-VRAM, or pure CPU.
-
Updated
Jul 15, 2026 - Python
DPython is a custom Python launcher for automatic single-GPU or multi-machine distributed training over LAN using Hugging Face Accelerate. It supports Windows, auto-detects local and remote GPUs, falls back safely to single-GPU when needed, and simplifies real-world multi-GPU training on low-VRAM systems.
-
Updated
Dec 27, 2025 - Python
Multi-GPU RAG pipeline for financial documents — parallel text and visual extraction over earnings transcripts, analyst notes and investor presentations.
-
Updated
Mar 6, 2026 - Python
Add this topic to your repo
To associate your repository with the multi-gpu topic, visit your repo's landing page and select "manage topics."