The Forge Cross-Platform Framework PC Windows, Steamdeck (native), Ray Tracing, macOS / iOS, Android, XBOX, PS4, PS5, Switch, Quest 2
-
Updated
Aug 27, 2026 - C++
The Forge Cross-Platform Framework PC Windows, Steamdeck (native), Ray Tracing, macOS / iOS, Android, XBOX, PS4, PS5, Switch, Quest 2
Toolkit for efficient experimentation with Speech Recognition, Text2Speech and NLP
Face recognition system for ID photos
Extract video features from raw videos using multiple GPUs. We support RAFT flow frames as well as S3D, I3D, R(2+1)D, VGGish, CLIP, and TIMM models.
<케라스 창시자에게 배우는 딥러닝 2판> 도서의 코드 저장소
Code for training py-faster-rcnn and py-R-FCN on multiple GPUs in caffe
Distributed tensors and Machine Learning framework with GPU and MPI acceleration in Python
The world's first CUDA implementation of Weakly-Compressible Smoothed Particle Hydrodynamics
A PyTorch implementation of the 'FaceNet' paper for training a facial recognition model with Triplet Loss using the glint360k dataset. A pre-trained model using Triplet Loss is available for download.
Multi-threaded GUI manager for mass creation of AI-generated art with support for multiple GPUs.
multi-gpu pre-training in one machine for BERT without horovod (Data Parallelism)
Package for writing high-level code for parallel high-performance stencil computations that can be deployed on both GPUs and CPUs
Efficient and Scalable Physics-Informed Deep Learning and Scientific Machine Learning on top of Tensorflow for multi-worker distributed computing
GPU-ready Dockerfile to run Stability.AI stable-diffusion model v2 with a simple web interface. Includes multi-GPUs support.
Kubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.
A dual-GPU DEM solver with complex grain geometry support
The Shamrock Framework, an open-source, multi-GPU hydrodynamics framework for astrophysics. Scales seamlessly from laptops to exascale supercomputers, supporting SPH, AMR, and more.
SSD-streaming inference engine for giant MoE models (Rust + CUDA). GLM 5.2 743B at 2 tok/s and Hy3 295B at 7 tok/s on two consumer 16GB GPUs. Zero-config multi-GPU: measures PCIe bandwidth, places attention and hot experts where they fit.
To associate your repository with the multi-gpu topic, visit your repo's landing page and select "manage topics."