Skip to content
View davetha's full-sized avatar

Sponsoring

@vllm-project

Block or report davetha

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Popular repositories Loading

  1. mi210-llm-stack mi210-llm-stack Public

    AMD MI210 (gfx90a/CDNA2) LLM inference optimization — TurboQuant, KIVI, per-layer KV types, FlashAttention, MoE expert caching

    Shell 11

  2. aiter-cdna2 aiter-cdna2 Public

    Run AMD AITER's hand-written ASM kernels on CDNA2 / gfx90a (MI210, MI250). 242 of 1,422 kernels translated, with the tests and benchmarks to prove which actually run.

    Python 10

  3. r9700-lru-expert-cache r9700-lru-expert-cache Public

    Device-side LRU expert cache + kernel-count patches: Qwen3.8-Flash-Next decode on 2x AMD R9700 (gfx1201), ROCm 10, tcclaviger vLLM fork

    Python 6 2

  4. vllm-expert-cache vllm-expert-cache Public

    vLLM plugin: device-side LRU/LFU expert cache for MoE offload. Keep N experts resident in VRAM, stream the rest from host RAM.

    Python 5 1

  5. mi210-vllm mi210-vllm Public

    Pinned, gate-verified vLLM deployment stack for AMD MI210 (gfx90a/CDNA2)

    Python 4

  6. cve-af-alg-block cve-af-alg-block Public

    Shell 3