Input-aware CUDA auto-scheduler for CSR SpMM, SDDMM, and sparse attention in graph neural networks.
-
Updated
Jul 22, 2026 - Python
Input-aware CUDA auto-scheduler for CSR SpMM, SDDMM, and sparse attention in graph neural networks.
Exploring computational scaling of spiking neural network simulations across hardware architectures
Official CUDA implementation of "Streaming Right Multiplication over Grammar-Compressed Matrices: A Memory-Bounded GPU Engine for Genotype and Graph Data" (ALENEX 2027). Enables memory-bounded matrix-vector products and graph algorithms on massive dataset via grammar compression.
CUDA/PyTorch operator for fused neighborhood sampling and mean aggregation, with reproducible KSEM 2026 experiments.
To associate your repository with the graph-processing topic, visit your repo's landing page and select "manage topics."