Repository navigation
v0.1.14
Highlights
- Reducer v2 (#2940, #3093, #3043, #3044, #3079, #3100):
T.alloc_reducerreworked into first-class deferred reduction epochs whose physical lowering is planned by layout inference (via a first-class PartialFragment layout). Adds loop-scoped epochs, conditional reducer finalization, and automatic vectorization of contiguous reducer updates. - Warp specialization schedules (#2892): new scheduling and materialization mechanism for warp-specialized kernels.
- Layout inference cost models (#2960, #3055, #3061): new IO-aware cost model for free-mode layout selection; register-count restored as the default, with an environment override to switch models.
- Unified backend resolution policy (#2318) plus backend split-up (#2855, #2870, #2850): backend selection is now resolved through a single policy, and builtin ops / Python op proxies are split per backend (CUDA/ROCm/Metal).
- Compilation speed: up to ~4x faster cold parallel/AOT compilation (#2809); Z3 solvers materialized lazily (#3105) and analyzer contexts isolated per kernel compilation (#2890).
- TMA rework: TMA copy lowering unified on CuTe algebra (#3106); TMA layouts made region-aware to keep slices contiguous (#3089).
Language
- Recycle
T.unroll(explicit=True)for early explicit unrolling (#2859) - Make the region bridge a builtin intrinsic (#2983)
- Expose
cluster_maskonT.tma_copy(#2932) - Unify contiguous stride construction under a single implementation (#3016); honor declared strides in pointer helpers (#3073)
- Stricter validation: reject symbolic
T.gemmtile dimensions with a clear message (#3113), validateT.gemmk_packarguments (#3094), reject non-positivearrive_countinalloc_barrier/alloc_cluster_barrier(#3112), rejectT.Parallelindexing of local buffers (#3041), rejectbreakin fully expanded loops (#3078)
CUDA
- tcgen05: pack logical TMEM buffers into shared
tcgen05.allocarenas (#2831); support half-subpartition (M=64) TMEM tiles intcgen05.ld/st(#2880); fix ld/st segment pointer advancement in b32 columns (#2952) - Select the widest legal WGMMA N instead of gcd (#2931)
- FP32x2 accumulation for reductions: per-reduce control (#3057) and a global PassConfig (#3128)
- Pre-SM80 fallback for bf16 atomic add (#2938);
int4x2/uint4x2codegen (#3036); 16-bit CUTLASS type overloads for fast-math,__ldg, and htan intrinsics (#3097, #3077, #3028, #2894) - Fixes: warp shuffle for half/bfloat16/FP8 (#3056), FP8 min/max codegen (#3047), UB in packed 8-bit vector stores (#3092), logical not for vectorized bool (#3117, also HIP), vectorized Select codegen (#2843), ldmatrix source offsets wrapped within shared-memory regions (#3110), NVRTC kernel handles isolated per adapter (#2950), flat CUDA include discovery for NVRTC (#2829), masked warpsync in in-warp allreduce (#2865)
- Remove
T.{reads,writes}forT.tma_{gather4,scatter4}(#3053); separate TMA atomic-add dtype support from layout encoding (#2846)
ROCm and other backends
- Remove the Composable Kernel dependency (#3111); ROCm CI re-enabled on a gfx942 runner (#2874, #2910)
- Fixes: preserve FP8 bits in warp shuffles (#3104), lower vector Select conditions lane-wise (#2889), emit a compiler barrier for
tl.sync_warpon HIP (#2872), reject sub-wavefront block sizes instead of crashing (#2918), resolve versioned device properties in the HIP stub (#2919); emit#linedirectives for the HIP target (#3058) - CPU backend: support atomic ops (#2941) and reduce ops (#2893)
- Metal: preserve pointer address spaces for byte offsets (#2925); resolve auto backend to torch and skip disk cache for torch (#2856)
- CuTeDSL: port backend intrinsics to CUTLASS DSL primitives (#2871)
Compiler / Transform
- Refactor the loop vectorization plan with ConstraintKind (#2935); always vectorize
T.Parallelloops (#3121); scalarize Select in automatic vectorization (#3060) - Add
VerifyBufferInit, a general buffer-initialization check (#2956) - Debug info: preserve source spans across lowering passes (#2966); emit
#linedirectives from TIR spans (#3048) - Fixes: don't drop syncs from the other if branch (#3085), fix wait parity for explicit mbarriers in pipelined loops (#3087), avoid int32 overflow in vector analysis (#3066), fix non-divisible nested modulo simplification (#3065), fix ties-away-from-zero round compile on bfloat16/float8 (#2873), fix absmax/abssum for uint dtypes (#2845), carry memory_order through vectorized atomic_add (#2924), fix unsigned zero-point decode underflow (#3118), keep cp.async operands in their address spaces (#2869), handle grid barriers and unbounded pointer ranges (#3050), deduplicate replicated reducer updates (#2881), bind symbolic coordinate ranges in FragmentThreadIndexProbe (#3096), reject non-round-tripping inferred layout inverses (#3090)
- Cleanup: remove obsolete compiler and runtime paths (#3086); remove the obsolete disable-fast-math pass config (#3098)
Runtime / JIT / Build
- Kernel cache: detect and repair corrupted cache entries (#3074), export libraries after disk-cache hits (#3116), remove the separate cache temporary directory (#3069)
- Allocate kernel outputs through the packed API (#2937); fix
get_parent_localsframe self-reference leak (#2934) - Support Cython 3.3 with the Python 3.9 limited API (#3068); fix CMake reconfigure aborting in FindPipCUDAToolkit before
project()(#3102)
Tooling / Ecosystem
- Official compile-only CLI (#3045)
- Unified pass instrumentation per compilation (#2923); Pass Visualizer driven by PassInstrument (#2866); show kernel name in the pass timing report (#2905)
- Data race check disabled by default, opt-in via env var (#2851)
- Open-source TileLang LSP announced (#2862); new agent skills: simplification (#3080), semantic validation (#3054), PR submission (#3082), backend architecture (#2900)
- Examples: generalize TCGEN05 GEMMs for Thor (sm110a) and add Stream-K scheduling (#2902); adopt multi-staged buffers in examples (#2836)
- Docs: add Sunrise-AI TANG (#3107), HYGON (#3101), and MetaX MACA (#3095) to supported platforms; link the multi-backend architecture design (#3123)
What's Changed
- [CUDA] Pack logical TMEM buffers into shared
tcgen05.allocarenas by @Rachmanino in #2831 - [Enhancement] Speed up cold parallel/AOT compilation up to ~4x by @cklxx in #2809
- [Docs] Refresh README news and onboarding by @LeiWang1999 in #2849
- [Enhancement] Disable data race check by default, opt-in via env var by @KellyFrog in #2851
- [BugFix] Fix absmax and abssum for uint dtypes by @jjppp in #2845
- [CUDA][TMA] Separate atomic-add dtype support from layout encoding by @LeiWang1999 in #2846
- [TIR][Python] Trim redundant op proxy wrappers by @SiriusNEO in #2850
- [CUDA][ROCm][Metal] Split backend-specific builtin ops by @LeiWang1999 in #2855
- [CUDA] Adopt multi-staged buffers in examples by @Yongqi-Zhuo in #2836
- [CI] [pre-commit.ci] autoupdate by @pre-commit-ci[bot] in #2861
- [Docs][LSP] Announce the open-source TileLang LSP by @LeiWang1999 in #2862
- [BugFix][Metal] Resolve auto backend to torch and skip disk cache for torch by @oraluben in #2856
- [Typo] Correct source spelling errors by @morluto in #2858
- [Refactor] Recycle
T.unroll(explicit=True)for early explicit unrolling by @Yongqi-Zhuo in #2859 - [BugFix] Use masked warpsync in in-warp allreduce by @jjppp in #2865
- [Debug][TIR] Drive Pass Visualizer with PassInstrument by @LeiWang1999 in #2866
- [Doc] Fix wrong loop bound in FlashAttention README example by @shanyi0228-web in #2868
- [Examples] Gate CUDA-only and flash_attn-dependent example tests by @andyluo7 in #2864
- [Testing] Gate CUDA-only tests so non-CUDA backends can run the suite by @andyluo7 in #2863
- [TIR][Python] Split backend-specific op proxies by @SiriusNEO in #2870
- [BugFix][ROCm] Emit a compiler barrier for tl.sync_warp on HIP by @andyluo7 in #2872
- [BugFix] Handle vectorized SelectNode in codegen_cuda by @jjppp in #2843
- [CI] Re-enable ROCm CI on a gfx942 runner by @andyluo7 in #2874
- [Cleanup] Replace root reproducers with CPU regression coverage by @GY-Bai in #2878
- [Compiler][Z3] Isolate analyzer contexts per kernel compilation by @LeiWang1999 in #2890
- [Backend] Add unified backend resolution policy by @SiriusNEO in #2318
- [Doc] ROCm CI is no longer disabled by @andyluo7 in #2896
- [BugFix] Deduplicate replicated reducer updates by @KellyFrog in #2881
- [Test] Add regression test for issue #2883 by @SiriusNEO in #2899
- [CPU] Support reduce ops on CPU by @penguin-wwy in #2893
- [Docs] Define backend architecture and integration skill by @SiriusNEO in #2900
- [Testing] Gate the issue #2883 regression test on CUDA by @andyluo7 in #2908
- [BugFix][ROCm] Lower vector Select conditions lane-wise by @morluto in #2889
- [CI] Use stable torch for the ROCm leg by @andyluo7 in #2910
- [Testing] Run portable regression tests on auto targets by @SiriusNEO in #2914
- [BugFix][Carver] Parse lettered SM arch strings in check_sm_version by @adityasingh2400 in #2891
- [Enhancement] Show kernel name in pass timing report by @penguin-wwy in #2905
- [CuTeDSL] Port backend intrinsics to CUTLASS DSL primitives by @cherichy in #2871
- [BugFix][ROCm] Reject sub-wavefront block sizes instead of crashing by @andyluo7 in #2918
- [Fix] Discover flat CUDA includes for NVRTC by @morluto in #2829
- [CI]: Bump pypa/cibuildwheel from 4.1 to 4.2 by @dependabot[bot] in #2930
- [ROCm] Resolve versioned device properties in HIP stub by @skyguan92 in #2919
- [BugFix] Skip DecoupleTypeCast on Evaluate roots to keep cp.async operands in their address spaces by @li-ruinan in #2869
- [Debug][TIR][JIT] Unify pass instrumentation per compilation by @LeiWang1999 in #2923
- [Cherry-Pick][BugFix] Fix get_parent_locals frame self-reference leak by @SiriusNEO in #2934
- [BugFix] Fix failed ties-away-from-zero round compile on bfloat16/float8 by @edragain2nd in #2873
- [JIT][FFI] Allocate kernel outputs through the packed API by @LeiWang1999 in #2937
- [Refactor][BugFix] Refactor the loop vectorization plan with ConstraintKind by @SiriusNEO in #2935
- [BugFix][CUDA] Provide htan overloads for fp16/bf16 tangent by @Ruihan11 in #2894
- [Layout] Support half-subpartition (M=64) TMEM tiles in tcgen05.ld/st by @Rachmanino in #2880
- [Feature] Warp specialization schedules and materialization by @Yongqi-Zhuo in #2892
- [CUDA] Add pre-SM80 fallback for bf16 atomic add by @Chennesxu in #2938
- [BugFix][Hopper] Select the widest legal WGMMA N instead of gcd by @bigSheep123 in #2931
- [Enhancement] Expose cluster_mask on T.tma_copy by @bigSheep123 in #2932
- [CPU] Support atomic ops on CPU by @penguin-wwy in #2941
- [Example] Generalize TCGEN05 GEMMs for Thor (sm110a) and add Stream-K scheduling by @xinhao-luo in #2902
- [Layout][CUDA] Support reinterpreting (dtype-changing) T.view aliases by @Yongqi-Zhuo in #2953
- [BugFix][CUDA] Advance tcgen05 ld/st segment pointers in b32 columns by @Yongqi-Zhuo in #2952
- [Testing] Pin the folded-base descriptor form for static ts slices by @Yongqi-Zhuo in #2951
- [Lang] Reducer v2: first-class deferred reduction epochs with planned physical lowering by @LeiWang1999 in #2940
- [CI][Examples] Remove TopK example from performance regression by @LeiWang1999 in #2958
- [BugFix][Metal] Preserve pointer address spaces for byte offsets by @GY-Bai in #2925
- [BugFix] Isolate NVRTC kernel handles per adapter by @SiriusNEO in #2950
- [Lang][TIR] Make region bridge a builtin intrinsic by @LeiWang1999 in #2983
- [Layout][Inference] Add IO-aware cost model for free-mode selection by @LeiWang1999 in #2960
- [Transform] Add VerifyBufferInit, a general buffer-initialization check by @RyanL2 in #2956
- [CUDA] Add __ldg overloads for 16-bit CUTLASS types by @Chennesxu in #3028
- [Analysis] Reject
T.Parallelindexing of local buffers by @SiriusNEO in #3041 - [Lang][Reducer] Support loop-scoped epochs and legacy default allocations by @LeiWang1999 in #3043
- [TIR][Transform] Preserve source spans across lowering passes by @penguin-wwy in #2966
- [Lang][Reducer] Allow conditional reducer finalization by @LeiWang1999 in #3044
- [CUDA] Fix FP8 min/max codegen by @Chennesxu in #3047
- [BugFix][Layout] Avoid thread-indexed wide reducer finalize readback by @LeiWang1999 in #3049
- [TIR][Transform] Handle grid barriers and unbounded pointer ranges by @LeiWang1999 in #3050
- [CodeGen] Emit #line directives from TIR spans by @penguin-wwy in #3048
- [Misc] Remove incorrect ASF license headers from src files by @penguin-wwy in #3051
- [Bugfix] Carry memory_order through vectorized atomic_add by @arcusbuilds in #2924
- [Layout][Inference] Restore register-count as the default cost model by @LeiWang1999 in #3055
- [CUDA] Remove
T.{reads,writes}forT.tma_{gather4,scatter4}by @Yongqi-Zhuo in #3053 - [Skill] Add TileLang semantic validation skill by @SiriusNEO in #3054
- [CUDA][Reduce] Add per-reduce control for FP32x2 accumulation by @LeiWang1999 in #3057
- [Layout][Config] Add environment override for layout cost model by @LeiWang1999 in #3061
- [CUDA] Fix warp shuffle for half, bfloat16, and FP8 by @Chennesxu in #3056
- [Build] Support Cython 3.3 with the Python 3.9 limited API by @LeiWang1999 in #3068
- [Runtime][Cache] Remove separate cache temporary directory by @LeiWang1999 in #3069
- [Language] Honor declared strides in pointer helpers by @zupengwang in #3073
- [BugFix] Scalarize Select in automatic vectorization by @KellyFrog in #3060
- [CodeGen][ROCm] Emit #line directives for HIP target by @penguin-wwy in #3058
- [Runtime][Cache] Detect and repair corrupted cache entries by @LeiWang1999 in #3074
- [Tool] Add official compile-only CLI by @LibertychaserUS in #3045
- [BugFix][CUDA] Support int4x2 and uint4x2 codegen by @SamJSui in #3036
- [CUDA] Bridge half-style math intrinsics for 16-bit CUTLASS types by @Chennesxu in #3077
- [Bugfix] Fix non-divisible nested modulo simplification by @haoyang9804 in #3065
- [Docs][CI] Add TileLang PR submission skill by @LeiWang1999 in #3082
- [Lang][Reducer] Vectorize contiguous reducer updates by @LeiWang1999 in #3079
- [SKILL] Add TileLang simplification skill by @SiriusNEO in #3080
- [Refactor] Remove obsolete compiler and runtime paths by @SiriusNEO in #3086
- [Language] Unify contiguous stride construction using a single implem… by @jjppp in #3016
- [Bugfix] Fix wait parity for explicit mbarriers in pipelined loops by @Yongqi-Zhuo in #3087
- [Bugfix] Make TMA layouts region-aware to keep slices contiguous by @Yongqi-Zhuo in #3089
- [Docs] Fix stale repository links by @morluto in #2857
- [Refactor] Give reducers a first-class PartialFragment layout solved by layout inference by @LeiWang1999 in #3093
- [BugFix] Bind symbolic coordinate ranges in FragmentThreadIndexProbe by @LeiWang1999 in #3096
- [Doc] Update support info with MetaX MACA backend by @Five-HZ in #3095
- [CUDA] Add 16-bit overloads for CUTLASS fast-math functions by @Chennesxu in #3097
- [Layout] Reject non-round-tripping inferred inverses by @KellyFrog in #3090
- [Bugfix] Validate T.gemm k_pack arguments by @WenzheWang in #3094
- [Cleanup] Remove obsolete disable-fast-math pass config by @SiriusNEO in #3098
- [Bugfix][CUDA] Fix UB in packed 8-bit vector stores by @yydhYYDH in #3092
- [BugFix][Transform] Avoid int32 overflow in vector analysis by @kobecai in #3066
- [Layout][Reducer] Preserve vectorized reducer update plans by @LeiWang1999 in #3100
- [Docs] List HYGON in README ecosystem hardware adapters by @warrenzzhou in #3101
- [Docs] Add Sunrise-AI TANG to platform support by @cratoroo in #3107
- [Bugfix][CMake] Fix reconfigure aborting in FindPipCUDAToolkit before project() by @SuperGoodGame in #3102
- [Compiler][Z3] Bump 3rdparty/tvm to materialize Z3 solvers lazily by @LeiWang1999 in #3105
- [ROCm] Preserve FP8 bits in warp shuffles by @andyluo7 in #3104
- [CUDA] Unify TMA copy lowering on CuTe algebra by @Yongqi-Zhuo in #3106
- [ROCm] Remove the Composable Kernel dependency by @LeiWang1999 in #3111
- [Lang] Reject non-positive arrive_count in alloc_barrier and alloc_cluster_barrier by @yurekami in #3112
- [Ci] Register TVM feature markers to get rid of PytestUnknownMarkWarning by @jjppp in #3115
- [Bugfix][Quantize] Fix unsigned zero-point decode underflow (#2947) by @Junius-Wynn in #3118
- [Bugfix] Reject symbolic T.gemm tile dimensions with a clear message by @jjppp in #3113
- [BugFix][CUDA][HIP] Handle logical not for vectorized bool by @jjppp in #3117
- [CUDA] Wrap ldmatrix source offsets within shared-memory regions by @Chennesxu in #3110
- [Runtime][Cache] Export libraries after disk-cache hits by @ZenAlexa in #3116
- [BugFix][Transform] Reject break in fully expanded loops (#3026) by @KellyFrog in #3078
- [BugFix] Always vectorize
T.Parallelloops by @Yongqi-Zhuo in #3121 - [BugFix] Don't drop syncs from the other if branch by @haoyang9804 in #3085
- [Docs] Link multi-backend architecture design by @SiriusNEO in #3123
- [Release] Bump version to 0.1.14 by @LeiWang1999 in #3119
- [CUDA][Reduce] Add PassConfig for FP32x2 accumulation by @SiriusNEO in #3128
New Contributors
- @shanyi0228-web made their first contribution in #2868
- @andyluo7 made their first contribution in #2864
- @adityasingh2400 made their first contribution in #2891
- @skyguan92 made their first contribution in #2919
- @edragain2nd made their first contribution in #2873
- @Ruihan11 made their first contribution in #2894
- @bigSheep123 made their first contribution in #2931
- @xinhao-luo made their first contribution in #2902
- @RyanL2 made their first contribution in #2956
- @arcusbuilds made their first contribution in #2924
- @zupengwang made their first contribution in #3073
- @LibertychaserUS made their first contribution in #3045
- @SamJSui made their first contribution in #3036
- @haoyang9804 made their first contribution in #3065
- @Five-HZ made their first contribution in #3095
- @WenzheWang made their first contribution in #3094
- @yydhYYDH made their first contribution in #3092
- @kobecai made their first contribution in #3066
- @warrenzzhou made their first contribution in #3101
- @cratoroo made their first contribution in #3107
- @SuperGoodGame made their first contribution in #3102
- @Junius-Wynn made their first contribution in #3118
- @ZenAlexa made their first contribution in #3116
Full Changelog: v0.1.13...v0.1.14