deepseek-ai / DeepGEMM
DeepGEMM: clean and efficient BLAS kernel library on GPU
See what the GitHub community is most excited about this week.
DeepGEMM: clean and efficient BLAS kernel library on GPU
FlashInfer: Kernel Library for LLM Serving
DeepEP: an efficient expert-parallel communication library
cuGraph - RAPIDS Graph Analytics Library
GPU accelerated decision optimization
Instant neural graphics primitives: lightning fast NeRF and more
LLM training in simple, raw C/CUDA
Mirage Persistent Kernel: Compiling LLMs into a MegaKernel