A no-cost infrastructure benchmark measuring the VRAM and throughput impact of NF4 (4-bit) quantization on LLMs
-
Updated
Dec 7, 2025 - Python
A no-cost infrastructure benchmark measuring the VRAM and throughput impact of NF4 (4-bit) quantization on LLMs
Enterprise-grade Power BI dashboard for tracking AI infrastructure unit economics. Features real-time GPU profitability ratios, idle waste detection (Zombie GPUs), and ESG carbon footprint reporting.
Archived history — active development moved to troycheng/cuda-kernel-optimizer
A massively parallel implementation of the Synthetic Aperture Radar (SAR) Time-Domain Backprojection algorithm optimized for NVIDIA GPUs using CUDA C++.
High-throughput, low-latency LLM inference platform for LLaMA-3 & Mistral — dynamic batching, KV-cache optimization, FP16/BF16 mixed precision, tensor parallelism, with PyTorch profiling, Prometheus/Grafana observability, Docker & Kubernetes (HPA) deployment.
"Fix for Docker Desktop 4.79.0 stuck on 'Starting the Docker Engine…' due to NVIDIA Container Toolkit conflicts in WSL2. No data loss, no downgrade needed."
GPU Stability Mastery 2026: Optimize Load, Fix Crashes, and Cut Heat
GPU Infrastructure Optimization
Optimized LSTM-based character-level text generator trained on Shakespeare, achieving 3.5x faster training with mixed precision.
Collection of Triton operators for transformer models.
Handwritten Flash Attention 2 CUDA kernel for Blackwell (SM120) with TMA, swizzle, double buffering & warp specialization
TrainSight: Sub-millisecond pre-flight LLM dataset profiler predicting HBM OOM crashes, padding waste, and attention entropy dispersion before GPU allocation. Built on the Two-Factor Law of LLM Compute.
Progressive CUDA SGEMM kernel(6) optimization (naive → vectorized) with cross-architecture profiling on RTX 3090 vs RTX 4050, benchmarked using Nsight Compute.
LLM inference engine implementing continuous batching, PagedAttention, prefix caching, prefill-decode disaggregation, and KV-cache-aware routing.
Fused Triton kernels for LIF spiking neuron dynamics. Bit-exact with snnTorch, up to 130x faster, 22% less memory.
GPU-accelerated multi-agent incident-triage system: LangGraph orchestration, TensorRT-LLM/Triton inference, Azure-native, fully observable.
Rule-evolving GPU kernel optimization loop — measurement feedback updates the rule table (matmul 6.4x, batched GEMM 4.5x on A100, KernelBench-validated)
Diagnose LLM inference cost and latency, then autotune serving configs with verified experiments. CLI + browser report viewer (v0.1), private autotune engine (v0.2), residual-driven search (v0.3).
DGX Spark (GB10/SM121) platform support for Meta's KernelAgent — auto-detect, hardware constraints, safe Triton configs
To associate your repository with the gpu-optimization topic, visit your repo's landing page and select "manage topics."