Toolkit for efficient experimentation with Speech Recognition, Text2Speech and NLP
-
Updated
May 11, 2021 - Python
Toolkit for efficient experimentation with Speech Recognition, Text2Speech and NLP
Extract video features from raw videos using multiple GPUs. We support RAFT flow frames as well as S3D, I3D, R(2+1)D, VGGish, CLIP, and TIMM models.
Multi-threaded GUI manager for mass creation of AI-generated art with support for multiple GPUs.
Face recognition system for ID photos
GPU-ready Dockerfile to run Stability.AI stable-diffusion model v2 with a simple web interface. Includes multi-GPUs support.
Distributed tensors and Machine Learning framework with GPU and MPI acceleration in Python
A PyTorch implementation of the 'FaceNet' paper for training a facial recognition model with Triplet Loss using the glint360k dataset. A pre-trained model using Triplet Loss is available for download.
TorchDR - PyTorch Dimensionality Reduction
Chains stable-diffusion-webui instances together to facilitate faster image generation.
multi-gpu pre-training in one machine for BERT without horovod (Data Parallelism)
Split FLUX.2 and LTX 2.3 across two GPUs (LAN or same-machine) — NVENC compresses activations live on the wire. Icarus (ComfyUI node) + Daedalus (back-half server).
Efficient and Scalable Physics-Informed Deep Learning and Scientific Machine Learning on top of Tensorflow for multi-worker distributed computing
Qwen3.8-Flash-Next on 2× RTX 3090: with 128 GB RAM up to 4,191 tok/s prefill · 111.5 tok/s decode (131K prompt), full 256K window at 2,865 tok/s; with 64 GB RAM 3,410 tok/s prefill · 84 tok/s decode. New: opt-in uncensored mode (runtime abliteration, no new weights).
Self-hosted AI-powered transcription platform with speaker diarization, search, and collaboration features. Built with Svelte, FastAPI, and Docker for easy deployment.
The LUMI AI Guide is designed to assist users in migrating their machine learning applications from smaller-scale computing environments to the LUMI supercomputer.
Ray-powered accelerator for MinerU, turning PDF → Markdown into a scalable, cluster-ready data infrastructure. 基于 Ray 的 MinerU 加速层,将 PDF → Markdown 构建为可扩展、面向集群的数据基础设施。
Neutron: A pytorch based implementation of Transformer and its variants.
The whole local AI stack in one executable: it runs and manages local AI models across every GPU, and it's a search engine you can talk to, with cited answers from your files, code, and the web. MCP server for coding agents, web crawler, TUI, CLI, REST API, Python library. No Ollama or LM Studio needed, works with both.
Unofficial reimplementation of DynaDUSt3R (Stereo4D, CVPR 2025): full training pipeline, sharded 4TB streaming data path, multi-GPU training. Weights and datasets released.
XReflection is a neat toolbox tailored for single-image reflection removal(SIRR). We offer state-of-the-art SIRR solutions for training and inference, with a high-performance data pipeline, multi-GPU/TPU/NPU support, and more!
To associate your repository with the multi-gpu topic, visit your repo's landing page and select "manage topics."