Repository navigation
Commit be668c6
Scope shared-memory bit-exact sizing to packed scalar NVFP4
The exact-bits shared storage sizing previously applied to every dtype
with element_bits % 8 != 0 (int4/uint4/uint2/...), halving e.g. int4
shared allocations while SharedByteOffsetToLogicalIndexOffset kept the
legacy bytes-per-element conversion for those dtypes — an inconsistent
size/index pair for any non-FP4 sub-byte buffer in a merged arena.
Restrict the new sizing to float4_e2m1fn scalar so it mirrors the
byte-offset -> logical-index special case exactly: NVFP4 buffers get
packed two-per-byte sizing and indexing, every other dtype keeps
bit-exact upstream behavior. This pass now changes nothing outside the
NVFP4 path.
Validated after rebuilding libtilelang.so: nvf4 language + example CLI +
quantize layout + access_ptr codegen + atom mma tests (129 passed),
example and WS benchmark --verify pass with a fresh JIT cache.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>1 parent 3e4984e commit be668c6
1 file changed
Lines changed: 18 additions & 7 deletions
File tree
- src/transform
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
83 | 83 | | |
84 | 84 | | |
85 | 85 | | |
| 86 | + | |
| 87 | + | |
| 88 | + | |
| 89 | + | |
| 90 | + | |
| 91 | + | |
| 92 | + | |
| 93 | + | |
| 94 | + | |
86 | 95 | | |
87 | | - | |
| 96 | + | |
| 97 | + | |
| 98 | + | |
| 99 | + | |
88 | 100 | | |
89 | 101 | | |
90 | 102 | | |
91 | 103 | | |
92 | 104 | | |
93 | 105 | | |
94 | | - | |
95 | | - | |
96 | | - | |
97 | | - | |
98 | | - | |
| 106 | + | |
| 107 | + | |
| 108 | + | |
| 109 | + | |
99 | 110 | | |
100 | 111 | | |
101 | 112 | | |
| |||
118 | 129 | | |
119 | 130 | | |
120 | 131 | | |
121 | | - | |
| 132 | + | |
122 | 133 | | |
123 | 134 | | |
124 | 135 | | |
| |||
0 commit comments