llama.cpp tuned for 2x Tesla P100 (sm_60). Decode 1.86x, 3-4x with MTP (up to ~70 t/s on code). Prefill 2x, full 262k context with vision. More numerically accurate than upstream.
-
Updated
Oct 2, 2026 - C++
llama.cpp tuned for 2x Tesla P100 (sm_60). Decode 1.86x, 3-4x with MTP (up to ~70 t/s on code). Prefill 2x, full 262k context with vision. More numerically accurate than upstream.
HRX, LOOM, Mac, and AMD running top speed kernels with a self optimizing engine.
A community effort to reduce neural-rendering GPU overhead while keeping enjoyable image quality. Includes a one-stop Wuthering Waves deployment helper.
Vulkan TBDR rendering optimizations for ARM Immortalis-G720 MC12 (Samsung Galaxy Tab S10 Ultra, Dimensity 9300+). Includes barrier reduction, memory stability fixes, missing feature fallbacks for Mali/Immortalis GPUs. ~60% FPS improvement in demanding titles.
Parallel Matrix Multiplication Optimization implementation using C++ and CUDA
GPU-accelerated image processing with automatic kernel tuning for AMD and NVIDIA GPUs.
Occupancy aware C++/CUDA operator fusion decision engine and physical hardware benchmark suite (Tesla T4 calibrated).
To associate your repository with the gpu-optimization topic, visit your repo's landing page and select "manage topics."