A General-purpose Task-parallel Programming System in C++
-
Updated
Sep 28, 2026 - C++
A General-purpose Task-parallel Programming System in C++
CUDA Core Compute Libraries
Sample codes for my CUDA programming book
Ecosystem of libraries and tools for writing and executing fast GPU code fully in Rust.
Safe rust wrapper around CUDA toolkit
Write Java. Run on GPUs. Fast.
A self-learning tutorail for CUDA High Performance Programing.
TinyChatEngine: On-Device LLM Inference Library
LLM notes, including model inference, transformer model structure, and llm framework code analysis notes.
Thin, unified, C++-flavored wrappers for the CUDA APIs
This is an archive of materials produced for an introductory class on CUDA programming at Stanford University in 2010
GPU Engineering for AI Systems
Adan: Adaptive Nesterov Momentum Algorithm for Faster Optimizing Deep Models
A native .NET LLM inference engine and agent runtime for GGUF models. TensorSharp provides a console application, a web-based chatbot interface, iPhone App, and Ollama/OpenAI-compatible HTTP APIs for programmatic access. It supports Windows/MacOS/iOS/Linux with full GPU capability
Static suckless single batch CUDA-only qwen3-0.6B mini inference engine
A simple GPU hash table implemented in CUDA using lock free techniques
A VisualBasic(.NET) language kernel and runtime for scientific data computing, deep learning, LLM inference, GPU acceleration, visualization and command-line data-science applications — running on .NET (net10.0) across Windows, Linux and macOS.
A curated list of best cuda programming books
An implementation of HIP that works on CPUs, across OSes.
To associate your repository with the cuda-programming topic, visit your repo's landing page and select "manage topics."