Repository navigation
Commit 6f75a92
committed
[Example] Add adaptive thread selection to sparse MLA backward
Default `bwd(..., threads=...)` to None and derive the launch width from
the head-block size instead of hard-coding 256.
This kernel is memory-pipe bound, so its throughput is set by how many
cp.async copies stay in flight, not by raw DRAM bandwidth. When the head
count is not split across blocks (block_H>=64), 256 threads (8 warps)
issue far more concurrent cp.async transfers than 128 (4 warps), greatly
raising in-flight bytes -- the main perf lever for this shape. Smaller
head blocks lack enough GEMM warp-tiling work to fill 256 threads, so
they fall back to 128 and still build.
For the deepseek_v32 shape (H=64, block_H=64) this resolves to 256,
matching the previous behavior.
Signed-off-by: Butterfingrz <13524387014@163.com>1 parent 3b37333 commit 6f75a92
1 file changed
Lines changed: 8 additions & 1 deletion
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
91 | 91 | | |
92 | 92 | | |
93 | 93 | | |
94 | | - | |
| 94 | + | |
95 | 95 | | |
96 | 96 | | |
97 | 97 | | |
| |||
122 | 122 | | |
123 | 123 | | |
124 | 124 | | |
| 125 | + | |
| 126 | + | |
| 127 | + | |
| 128 | + | |
| 129 | + | |
| 130 | + | |
| 131 | + | |
125 | 132 | | |
126 | 133 | | |
127 | 134 | | |
| |||
0 commit comments