You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Official implementation of CoCoFold2: scalable latent refinement of protein structure predictions from limited-particle cryo-EM data using a frozen Protenix-v1 diffusion prior. Supports single-GPU and component-parallel refinement.
Genuine multi-GPU FSDP full fine-tune of a 7B across 4x A100 (closes the scale gap), with the tied-embeddings, collective-save, and checkpoint-consolidation gotchas made concrete.
Self-hosted private AI workspace on owned Ubuntu hardware. llama.cpp layer-splits Qwen3.8-27B across an RTX 3060 and RTX 3070; a FastAPI app adds login-gated streaming chat, SearXNG search, page fetch, weather, Hugging Face and optional document retrieval. No LAN or public listener - remote access via Cloudflare Tunnel. AGPL-3.0.
Simulated Multi-GPU inference engine implementing Megatron-style Tensor Parallelism, GPipe Pipeline Parallelism, and KV Cache Sharding from scratch. Features bit-identical fidelity validation on Qwen2-0.5B weights and analytical communication cost modeling.
Quality-first local MiniMax H3 video-series generation and loopback API for dual RTX 4090 workstations, with native audio, references, P8/P9 continuity, and preserved artifacts.
Lightweight terminal launcher and auto-optimizer for llm models using llama.cpp with hardware detection, tensor sharding, benchmarks, presets, context tuning, and OpenAI-compatible serving.
Experiments from 'NLP with Transformers' (fine-tuning, summarization, knowledge distillation) as standalone scripts, plus multi-GPU training with Hugging Face Accelerate: DDP gives ~1.8x the throughput of DataParallel on 2x RTX 4060 Ti