Skip to content

v0.1.14

Choose a tag to compare

@LeiWang1999 LeiWang1999 released this 02 Sep 05:33
· 91 commits to main since this release
7e3bbf3

Highlights

  • Reducer v2 (#2940, #3093, #3043, #3044, #3079, #3100): T.alloc_reducer reworked into first-class deferred reduction epochs whose physical lowering is planned by layout inference (via a first-class PartialFragment layout). Adds loop-scoped epochs, conditional reducer finalization, and automatic vectorization of contiguous reducer updates.
  • Warp specialization schedules (#2892): new scheduling and materialization mechanism for warp-specialized kernels.
  • Layout inference cost models (#2960, #3055, #3061): new IO-aware cost model for free-mode layout selection; register-count restored as the default, with an environment override to switch models.
  • Unified backend resolution policy (#2318) plus backend split-up (#2855, #2870, #2850): backend selection is now resolved through a single policy, and builtin ops / Python op proxies are split per backend (CUDA/ROCm/Metal).
  • Compilation speed: up to ~4x faster cold parallel/AOT compilation (#2809); Z3 solvers materialized lazily (#3105) and analyzer contexts isolated per kernel compilation (#2890).
  • TMA rework: TMA copy lowering unified on CuTe algebra (#3106); TMA layouts made region-aware to keep slices contiguous (#3089).

Language

  • Recycle T.unroll(explicit=True) for early explicit unrolling (#2859)
  • Make the region bridge a builtin intrinsic (#2983)
  • Expose cluster_mask on T.tma_copy (#2932)
  • Unify contiguous stride construction under a single implementation (#3016); honor declared strides in pointer helpers (#3073)
  • Stricter validation: reject symbolic T.gemm tile dimensions with a clear message (#3113), validate T.gemm k_pack arguments (#3094), reject non-positive arrive_count in alloc_barrier/alloc_cluster_barrier (#3112), reject T.Parallel indexing of local buffers (#3041), reject break in fully expanded loops (#3078)

CUDA

  • tcgen05: pack logical TMEM buffers into shared tcgen05.alloc arenas (#2831); support half-subpartition (M=64) TMEM tiles in tcgen05.ld/st (#2880); fix ld/st segment pointer advancement in b32 columns (#2952)
  • Select the widest legal WGMMA N instead of gcd (#2931)
  • FP32x2 accumulation for reductions: per-reduce control (#3057) and a global PassConfig (#3128)
  • Pre-SM80 fallback for bf16 atomic add (#2938); int4x2/uint4x2 codegen (#3036); 16-bit CUTLASS type overloads for fast-math, __ldg, and htan intrinsics (#3097, #3077, #3028, #2894)
  • Fixes: warp shuffle for half/bfloat16/FP8 (#3056), FP8 min/max codegen (#3047), UB in packed 8-bit vector stores (#3092), logical not for vectorized bool (#3117, also HIP), vectorized Select codegen (#2843), ldmatrix source offsets wrapped within shared-memory regions (#3110), NVRTC kernel handles isolated per adapter (#2950), flat CUDA include discovery for NVRTC (#2829), masked warpsync in in-warp allreduce (#2865)
  • Remove T.{reads,writes} for T.tma_{gather4,scatter4} (#3053); separate TMA atomic-add dtype support from layout encoding (#2846)

ROCm and other backends

  • Remove the Composable Kernel dependency (#3111); ROCm CI re-enabled on a gfx942 runner (#2874, #2910)
  • Fixes: preserve FP8 bits in warp shuffles (#3104), lower vector Select conditions lane-wise (#2889), emit a compiler barrier for tl.sync_warp on HIP (#2872), reject sub-wavefront block sizes instead of crashing (#2918), resolve versioned device properties in the HIP stub (#2919); emit #line directives for the HIP target (#3058)
  • CPU backend: support atomic ops (#2941) and reduce ops (#2893)
  • Metal: preserve pointer address spaces for byte offsets (#2925); resolve auto backend to torch and skip disk cache for torch (#2856)
  • CuTeDSL: port backend intrinsics to CUTLASS DSL primitives (#2871)

Compiler / Transform

  • Refactor the loop vectorization plan with ConstraintKind (#2935); always vectorize T.Parallel loops (#3121); scalarize Select in automatic vectorization (#3060)
  • Add VerifyBufferInit, a general buffer-initialization check (#2956)
  • Debug info: preserve source spans across lowering passes (#2966); emit #line directives from TIR spans (#3048)
  • Fixes: don't drop syncs from the other if branch (#3085), fix wait parity for explicit mbarriers in pipelined loops (#3087), avoid int32 overflow in vector analysis (#3066), fix non-divisible nested modulo simplification (#3065), fix ties-away-from-zero round compile on bfloat16/float8 (#2873), fix absmax/abssum for uint dtypes (#2845), carry memory_order through vectorized atomic_add (#2924), fix unsigned zero-point decode underflow (#3118), keep cp.async operands in their address spaces (#2869), handle grid barriers and unbounded pointer ranges (#3050), deduplicate replicated reducer updates (#2881), bind symbolic coordinate ranges in FragmentThreadIndexProbe (#3096), reject non-round-tripping inferred layout inverses (#3090)
  • Cleanup: remove obsolete compiler and runtime paths (#3086); remove the obsolete disable-fast-math pass config (#3098)

Runtime / JIT / Build

  • Kernel cache: detect and repair corrupted cache entries (#3074), export libraries after disk-cache hits (#3116), remove the separate cache temporary directory (#3069)
  • Allocate kernel outputs through the packed API (#2937); fix get_parent_locals frame self-reference leak (#2934)
  • Support Cython 3.3 with the Python 3.9 limited API (#3068); fix CMake reconfigure aborting in FindPipCUDAToolkit before project() (#3102)

Tooling / Ecosystem

  • Official compile-only CLI (#3045)
  • Unified pass instrumentation per compilation (#2923); Pass Visualizer driven by PassInstrument (#2866); show kernel name in the pass timing report (#2905)
  • Data race check disabled by default, opt-in via env var (#2851)
  • Open-source TileLang LSP announced (#2862); new agent skills: simplification (#3080), semantic validation (#3054), PR submission (#3082), backend architecture (#2900)
  • Examples: generalize TCGEN05 GEMMs for Thor (sm110a) and add Stream-K scheduling (#2902); adopt multi-staged buffers in examples (#2836)
  • Docs: add Sunrise-AI TANG (#3107), HYGON (#3101), and MetaX MACA (#3095) to supported platforms; link the multi-backend architecture design (#3123)

What's Changed

  • [CUDA] Pack logical TMEM buffers into shared tcgen05.alloc arenas by @Rachmanino in #2831
  • [Enhancement] Speed up cold parallel/AOT compilation up to ~4x by @cklxx in #2809
  • [Docs] Refresh README news and onboarding by @LeiWang1999 in #2849
  • [Enhancement] Disable data race check by default, opt-in via env var by @KellyFrog in #2851
  • [BugFix] Fix absmax and abssum for uint dtypes by @jjppp in #2845
  • [CUDA][TMA] Separate atomic-add dtype support from layout encoding by @LeiWang1999 in #2846
  • [TIR][Python] Trim redundant op proxy wrappers by @SiriusNEO in #2850
  • [CUDA][ROCm][Metal] Split backend-specific builtin ops by @LeiWang1999 in #2855
  • [CUDA] Adopt multi-staged buffers in examples by @Yongqi-Zhuo in #2836
  • [CI] [pre-commit.ci] autoupdate by @pre-commit-ci[bot] in #2861
  • [Docs][LSP] Announce the open-source TileLang LSP by @LeiWang1999 in #2862
  • [BugFix][Metal] Resolve auto backend to torch and skip disk cache for torch by @oraluben in #2856
  • [Typo] Correct source spelling errors by @morluto in #2858
  • [Refactor] Recycle T.unroll(explicit=True) for early explicit unrolling by @Yongqi-Zhuo in #2859
  • [BugFix] Use masked warpsync in in-warp allreduce by @jjppp in #2865
  • [Debug][TIR] Drive Pass Visualizer with PassInstrument by @LeiWang1999 in #2866
  • [Doc] Fix wrong loop bound in FlashAttention README example by @shanyi0228-web in #2868
  • [Examples] Gate CUDA-only and flash_attn-dependent example tests by @andyluo7 in #2864
  • [Testing] Gate CUDA-only tests so non-CUDA backends can run the suite by @andyluo7 in #2863
  • [TIR][Python] Split backend-specific op proxies by @SiriusNEO in #2870
  • [BugFix][ROCm] Emit a compiler barrier for tl.sync_warp on HIP by @andyluo7 in #2872
  • [BugFix] Handle vectorized SelectNode in codegen_cuda by @jjppp in #2843
  • [CI] Re-enable ROCm CI on a gfx942 runner by @andyluo7 in #2874
  • [Cleanup] Replace root reproducers with CPU regression coverage by @GY-Bai in #2878
  • [Compiler][Z3] Isolate analyzer contexts per kernel compilation by @LeiWang1999 in #2890
  • [Backend] Add unified backend resolution policy by @SiriusNEO in #2318
  • [Doc] ROCm CI is no longer disabled by @andyluo7 in #2896
  • [BugFix] Deduplicate replicated reducer updates by @KellyFrog in #2881
  • [Test] Add regression test for issue #2883 by @SiriusNEO in #2899
  • [CPU] Support reduce ops on CPU by @penguin-wwy in #2893
  • [Docs] Define backend architecture and integration skill by @SiriusNEO in #2900
  • [Testing] Gate the issue #2883 regression test on CUDA by @andyluo7 in #2908
  • [BugFix][ROCm] Lower vector Select conditions lane-wise by @morluto in #2889
  • [CI] Use stable torch for the ROCm leg by @andyluo7 in #2910
  • [Testing] Run portable regression tests on auto targets by @SiriusNEO in #2914
  • [BugFix][Carver] Parse lettered SM arch strings in check_sm_version by @adityasingh2400 in #2891
  • [Enhancement] Show kernel name in pass timing report by @penguin-wwy in #2905
  • [CuTeDSL] Port backend intrinsics to CUTLASS DSL primitives by @cherichy in #2871
  • [BugFix][ROCm] Reject sub-wavefront block sizes instead of crashing by @andyluo7 in #2918
  • [Fix] Discover flat CUDA includes for NVRTC by @morluto in #2829
  • [CI]: Bump pypa/cibuildwheel from 4.1 to 4.2 by @dependabot[bot] in #2930
  • [ROCm] Resolve versioned device properties in HIP stub by @skyguan92 in #2919
  • [BugFix] Skip DecoupleTypeCast on Evaluate roots to keep cp.async operands in their address spaces by @li-ruinan in #2869
  • [Debug][TIR][JIT] Unify pass instrumentation per compilation by @LeiWang1999 in #2923
  • [Cherry-Pick][BugFix] Fix get_parent_locals frame self-reference leak by @SiriusNEO in #2934
  • [BugFix] Fix failed ties-away-from-zero round compile on bfloat16/float8 by @edragain2nd in #2873
  • [JIT][FFI] Allocate kernel outputs through the packed API by @LeiWang1999 in #2937
  • [Refactor][BugFix] Refactor the loop vectorization plan with ConstraintKind by @SiriusNEO in #2935
  • [BugFix][CUDA] Provide htan overloads for fp16/bf16 tangent by @Ruihan11 in #2894
  • [Layout] Support half-subpartition (M=64) TMEM tiles in tcgen05.ld/st by @Rachmanino in #2880
  • [Feature] Warp specialization schedules and materialization by @Yongqi-Zhuo in #2892
  • [CUDA] Add pre-SM80 fallback for bf16 atomic add by @Chennesxu in #2938
  • [BugFix][Hopper] Select the widest legal WGMMA N instead of gcd by @bigSheep123 in #2931
  • [Enhancement] Expose cluster_mask on T.tma_copy by @bigSheep123 in #2932
  • [CPU] Support atomic ops on CPU by @penguin-wwy in #2941
  • [Example] Generalize TCGEN05 GEMMs for Thor (sm110a) and add Stream-K scheduling by @xinhao-luo in #2902
  • [Layout][CUDA] Support reinterpreting (dtype-changing) T.view aliases by @Yongqi-Zhuo in #2953
  • [BugFix][CUDA] Advance tcgen05 ld/st segment pointers in b32 columns by @Yongqi-Zhuo in #2952
  • [Testing] Pin the folded-base descriptor form for static ts slices by @Yongqi-Zhuo in #2951
  • [Lang] Reducer v2: first-class deferred reduction epochs with planned physical lowering by @LeiWang1999 in #2940
  • [CI][Examples] Remove TopK example from performance regression by @LeiWang1999 in #2958
  • [BugFix][Metal] Preserve pointer address spaces for byte offsets by @GY-Bai in #2925
  • [BugFix] Isolate NVRTC kernel handles per adapter by @SiriusNEO in #2950
  • [Lang][TIR] Make region bridge a builtin intrinsic by @LeiWang1999 in #2983
  • [Layout][Inference] Add IO-aware cost model for free-mode selection by @LeiWang1999 in #2960
  • [Transform] Add VerifyBufferInit, a general buffer-initialization check by @RyanL2 in #2956
  • [CUDA] Add __ldg overloads for 16-bit CUTLASS types by @Chennesxu in #3028
  • [Analysis] Reject T.Parallel indexing of local buffers by @SiriusNEO in #3041
  • [Lang][Reducer] Support loop-scoped epochs and legacy default allocations by @LeiWang1999 in #3043
  • [TIR][Transform] Preserve source spans across lowering passes by @penguin-wwy in #2966
  • [Lang][Reducer] Allow conditional reducer finalization by @LeiWang1999 in #3044
  • [CUDA] Fix FP8 min/max codegen by @Chennesxu in #3047
  • [BugFix][Layout] Avoid thread-indexed wide reducer finalize readback by @LeiWang1999 in #3049
  • [TIR][Transform] Handle grid barriers and unbounded pointer ranges by @LeiWang1999 in #3050
  • [CodeGen] Emit #line directives from TIR spans by @penguin-wwy in #3048
  • [Misc] Remove incorrect ASF license headers from src files by @penguin-wwy in #3051
  • [Bugfix] Carry memory_order through vectorized atomic_add by @arcusbuilds in #2924
  • [Layout][Inference] Restore register-count as the default cost model by @LeiWang1999 in #3055
  • [CUDA] Remove T.{reads,writes} for T.tma_{gather4,scatter4} by @Yongqi-Zhuo in #3053
  • [Skill] Add TileLang semantic validation skill by @SiriusNEO in #3054
  • [CUDA][Reduce] Add per-reduce control for FP32x2 accumulation by @LeiWang1999 in #3057
  • [Layout][Config] Add environment override for layout cost model by @LeiWang1999 in #3061
  • [CUDA] Fix warp shuffle for half, bfloat16, and FP8 by @Chennesxu in #3056
  • [Build] Support Cython 3.3 with the Python 3.9 limited API by @LeiWang1999 in #3068
  • [Runtime][Cache] Remove separate cache temporary directory by @LeiWang1999 in #3069
  • [Language] Honor declared strides in pointer helpers by @zupengwang in #3073
  • [BugFix] Scalarize Select in automatic vectorization by @KellyFrog in #3060
  • [CodeGen][ROCm] Emit #line directives for HIP target by @penguin-wwy in #3058
  • [Runtime][Cache] Detect and repair corrupted cache entries by @LeiWang1999 in #3074
  • [Tool] Add official compile-only CLI by @LibertychaserUS in #3045
  • [BugFix][CUDA] Support int4x2 and uint4x2 codegen by @SamJSui in #3036
  • [CUDA] Bridge half-style math intrinsics for 16-bit CUTLASS types by @Chennesxu in #3077
  • [Bugfix] Fix non-divisible nested modulo simplification by @haoyang9804 in #3065
  • [Docs][CI] Add TileLang PR submission skill by @LeiWang1999 in #3082
  • [Lang][Reducer] Vectorize contiguous reducer updates by @LeiWang1999 in #3079
  • [SKILL] Add TileLang simplification skill by @SiriusNEO in #3080
  • [Refactor] Remove obsolete compiler and runtime paths by @SiriusNEO in #3086
  • [Language] Unify contiguous stride construction using a single implem… by @jjppp in #3016
  • [Bugfix] Fix wait parity for explicit mbarriers in pipelined loops by @Yongqi-Zhuo in #3087
  • [Bugfix] Make TMA layouts region-aware to keep slices contiguous by @Yongqi-Zhuo in #3089
  • [Docs] Fix stale repository links by @morluto in #2857
  • [Refactor] Give reducers a first-class PartialFragment layout solved by layout inference by @LeiWang1999 in #3093
  • [BugFix] Bind symbolic coordinate ranges in FragmentThreadIndexProbe by @LeiWang1999 in #3096
  • [Doc] Update support info with MetaX MACA backend by @Five-HZ in #3095
  • [CUDA] Add 16-bit overloads for CUTLASS fast-math functions by @Chennesxu in #3097
  • [Layout] Reject non-round-tripping inferred inverses by @KellyFrog in #3090
  • [Bugfix] Validate T.gemm k_pack arguments by @WenzheWang in #3094
  • [Cleanup] Remove obsolete disable-fast-math pass config by @SiriusNEO in #3098
  • [Bugfix][CUDA] Fix UB in packed 8-bit vector stores by @yydhYYDH in #3092
  • [BugFix][Transform] Avoid int32 overflow in vector analysis by @kobecai in #3066
  • [Layout][Reducer] Preserve vectorized reducer update plans by @LeiWang1999 in #3100
  • [Docs] List HYGON in README ecosystem hardware adapters by @warrenzzhou in #3101
  • [Docs] Add Sunrise-AI TANG to platform support by @cratoroo in #3107
  • [Bugfix][CMake] Fix reconfigure aborting in FindPipCUDAToolkit before project() by @SuperGoodGame in #3102
  • [Compiler][Z3] Bump 3rdparty/tvm to materialize Z3 solvers lazily by @LeiWang1999 in #3105
  • [ROCm] Preserve FP8 bits in warp shuffles by @andyluo7 in #3104
  • [CUDA] Unify TMA copy lowering on CuTe algebra by @Yongqi-Zhuo in #3106
  • [ROCm] Remove the Composable Kernel dependency by @LeiWang1999 in #3111
  • [Lang] Reject non-positive arrive_count in alloc_barrier and alloc_cluster_barrier by @yurekami in #3112
  • [Ci] Register TVM feature markers to get rid of PytestUnknownMarkWarning by @jjppp in #3115
  • [Bugfix][Quantize] Fix unsigned zero-point decode underflow (#2947) by @Junius-Wynn in #3118
  • [Bugfix] Reject symbolic T.gemm tile dimensions with a clear message by @jjppp in #3113
  • [BugFix][CUDA][HIP] Handle logical not for vectorized bool by @jjppp in #3117
  • [CUDA] Wrap ldmatrix source offsets within shared-memory regions by @Chennesxu in #3110
  • [Runtime][Cache] Export libraries after disk-cache hits by @ZenAlexa in #3116
  • [BugFix][Transform] Reject break in fully expanded loops (#3026) by @KellyFrog in #3078
  • [BugFix] Always vectorize T.Parallel loops by @Yongqi-Zhuo in #3121
  • [BugFix] Don't drop syncs from the other if branch by @haoyang9804 in #3085
  • [Docs] Link multi-backend architecture design by @SiriusNEO in #3123
  • [Release] Bump version to 0.1.14 by @LeiWang1999 in #3119
  • [CUDA][Reduce] Add PassConfig for FP32x2 accumulation by @SiriusNEO in #3128

New Contributors

Full Changelog: v0.1.13...v0.1.14