- Houston, Texas
Popular repositories Loading
-
mi210-llm-stack
mi210-llm-stack PublicAMD MI210 (gfx90a/CDNA2) LLM inference optimization — TurboQuant, KIVI, per-layer KV types, FlashAttention, MoE expert caching
Shell 11
-
aiter-cdna2
aiter-cdna2 PublicRun AMD AITER's hand-written ASM kernels on CDNA2 / gfx90a (MI210, MI250). 242 of 1,422 kernels translated, with the tests and benchmarks to prove which actually run.
Python 10
-
r9700-lru-expert-cache
r9700-lru-expert-cache PublicDevice-side LRU expert cache + kernel-count patches: Qwen3.8-Flash-Next decode on 2x AMD R9700 (gfx1201), ROCm 10, tcclaviger vLLM fork
-
vllm-expert-cache
vllm-expert-cache PublicvLLM plugin: device-side LRU/LFU expert cache for MoE offload. Keep N experts resident in VRAM, stream the rest from host RAM.
-
mi210-vllm
mi210-vllm PublicPinned, gate-verified vLLM deployment stack for AMD MI210 (gfx90a/CDNA2)
Python 4
-
If the problem persists, check the GitHub status page or contact support.





