An Open-source book with Modern CUDA learning Notes for Beginners - BF16/FP8/FP4, HGEMM, CuTe, Flash-Attention, etc.
-
Updated
Oct 7, 2026 - Cuda
An Open-source book with Modern CUDA learning Notes for Beginners - BF16/FP8/FP4, HGEMM, CuTe, Flash-Attention, etc.
Kernl lets you run PyTorch transformer models several times faster on GPU with a single line of code, and is designed to be easily hackable.
row-major matmul optimization
Fast, differentiable sorting and ranking in PyTorch
GPU 性能与 AI Infra 学习项目:CUDA/Triton 算子、NCU/NSYS、vLLM/SGLang/TRT-LLM/ms-swift、PyTorch/DeepSpeed/ms-swift 训练、并行架构
A performance comparison of standard matrix functions between CPU and GPU using Nvidia CUDA on Visual Studio using C++
a custom CUDA kernel for windowed matrix multiplication
SNU CSE Scalable High Performance Computing (M1522.006700) - 2023 Autumn
A production-ready, high-performance CUDA Runtime and Driver API library for Zig.
From-scratch SM80 (A100/A800) CUDA kernels for GDN (Gated DeltaNet) and QSA (query-key sparse attention). Apache-2.0.
pointwise silu(gate) * up → Y cuda kernel
The Road to be CUDA Expert
NYCU 2025 Spring Computer Architecture(CA) 陽明交大 劉志尉 計算機結構
Arbitrary Precision Prime Number Finder
Progressive CUDA GEMM optimization from naive to warp-tiled — benchmarked against cuBLAS on NVIDIA T4
CUDA and Futhark implementations of core GPU / data-parallel techniques: shared & constant memory, coalescing, thread synchronization, list homomorphisms, and sparse matrix-vector products.
Custom PyTorch CUDA kernel implementing optimized ReLU activation with vectorization, performance profiling, and memory analysis on Tesla T4 GPU achieving 75% bandwidth efficiency.
To associate your repository with the cuda-kernel topic, visit your repo's landing page and select "manage topics."