Skip to content

v0.1.6.post2

Choose a tag to compare

@LeiWang1999 LeiWang1999 released this 31 Oct 01:00
· 1206 commits to main since this release
c37621c

The Last Release for Python 3.8 (without tvm-ffi) 馃殌

What's Changed

  • [Analyzer] Enhance ConstIntBoundAnalyzer and IntervalSet with modular set analysis by @LeiWang1999 in #856
  • [Doc] Optimize the quickstart guide for clarity and not just for CUDA by @LeiWang1999 in #858
  • [TMA] Bugfix when a shared buffer is both issued with tma store and tma load by @LeiWang1999 in #857
  • [AMD][MLA] Fix mla autotune for rocm by @LeiWang1999 in #861
  • [Bugfix] Ensure correct handling for cases where seq_q<seq_kv in flash attention examples by @Rachmanino in #864
  • [AMD] refactor MatrixCoreIntrinEmitter by @Paran0idy in #860
  • [Feat] Add fast sine and cosine definitions in CUDA templates by @Rachmanino in #865
  • [Layout] Support layout forward with multi dimension by @LeiWang1999 in #867
  • [Autotune][Conv] optimize convolution examples to use autotune by @LeiWang1999 in #866
  • [Example] Add examples to support efficient attention sink forward process by @Rachmanino in #853
  • [Parser] Adapt Parser to work with Python 3.8 in some cases by @LeiWang1999 in #869
  • [Fix] Fix bug 0905: tilelang doesn't vectorize B[i,j] = c[i] + A[i,j] by @kurisu6912 in #798
  • [Language] Support sequence comparisons by @LeiWang1999 in #872
  • [Language] Support loop_break primitive by @chengyupku in #873
  • [Bugfix] Use ExprDeepEqual instead of StructuralEqual when merge consecutive If stmt by @LeiWang1999 in #876
  • [Language] Support atomic add with ret by @LeiWang1999 in #870
  • [Cython] Remove an incorrect check by @LJC00118 in #880
  • Update amd_ci.yml by @Alex4210987 in #881
  • [FastMath] Disable default TVM fastmath intrinsic dispatch and add explicit fastmath op to invoke by @LeiWang1999 in #875
  • [Example] Add efficient attention sink backward implementations and tests by @Rachmanino in #877
  • [Precision] Introduce T.ieee_rsqrt and related high precision op by @LeiWang1999 in #882
  • [Dist] Provide an option to include commit ID in version by @LeiWang1999 in #884
  • [Example] Optimize sink attention forward via swizzled layout and report benchmark results by @Rachmanino in #885
  • [Layout] Introduce Flexible Parallel to Support T.serial and local buffers inside T.Parallel loop by @LeiWang1999 in #844
  • [Bugfix][Enhancement] Fix a bug in previous commit and enhance cuda backend by @Hamerlate in #887
  • [Bugfix] Fix CopyNode Lower method to include disable_tma flag in GetCopyInst by @Rachmanino in #888
  • [Layout] Fix plot layout by @Paran0idy in #890
  • [Example] Add example by @LeiWang1999 in #894
  • [News] Add announcement of support for Huawei Ascend chips by @xwhzz in #895
  • [Example] Add sparse mla examples by @LeiWang1999 in #896
  • [Typo] Fix backend name for Huawei Ascend by @xwhzz in #898
  • [CI] Legalize math related test by @LeiWang1999 in #899
  • [Bugfix] Fix flops comp and softmax scale in mla by @Edenzzzz in #900
  • [Example] Specify a fixed commit for the flash-linear-attention repository and optimize nsa examples by @LeiWang1999 in #913
  • [CI] optimize CI time for sparse gemm by @botbw in #906
  • [Enhancement] Include compile flags into the hash key of cached kernels by @Rachmanino in #911
  • [Bugfix] Fix saving kernel source code where JITKernel.artifact is None by @zjudmd1015 in #921
  • [CI] Refactor import paths in dequantization examples to use dequantize_utils by @LeiWang1999 in #914
  • [Example] Add MLA decode ws example by @chengyupku in #928
  • [CI] Fix documentation runner by adding 'nvidia' tag by @xwhzz in #927
  • [Layout] Strict annotate completed replicated layout for fragment with constant index by @LeiWang1999 in #929
  • [Bugfix] Fix tensor memory copy layout by @Hamerlate in #933
  • [Example] Optimize online_softmax example by @lijinpei in #934
  • [Example] Add correctness assert into dsa example by @LeiWang1999 in #937
  • [Enhancement] Enhance and add new GQA backward examples for Hopper by @Rachmanino in #930
  • [Enhancement] Fix lint to improve grouped GEMM performance with TMA by @Cunxiao2002 in #938
  • [Example] Introduce split+sum template, and optimize atomic_add performance for bwd examples by @LeiWang1999 in #940
  • [Example] Disable TMA and enable FastMath for NSA Examples (#941) by @LeiWang1999 in #941
  • [Example] Revert the atomic/split&sum templates in MHA backward examples by @Rachmanino in #943
  • [Example] Add sparse mla bwd example for deepseek_v32 by @Zhichenzzz in #919
  • [Profiler]Adds CUPTI profiler support by @Cunxiao2002 in #936
  • [Enhancement] Support Copy for Buffer Load witih scalar indices by @LeiWang1999 in #946
  • [Code Style] Refine nvrtc compile related check style by @BBuf in #945
  • [Backend] Add metal backend by @oraluben in #799
  • [CI] enable dependabot for GHA workflows by @XuehaiPan in #950
  • Modify the SM architecture number to support Thor鈥檚 sm110. by @iloveai8086 in #957
  • [CI] auto-cancel in-progress PR CI when new commits are pushed by @XuehaiPan in #956
  • [bug] fix type object is not subscriptable in py38 by @BBuf in #959
  • [Bugfix][Doc] Add astroid version constraint to requirements.txt by @xwhzz in #958
  • [CI]: Bump actions/setup-python from 2 to 6 by @dependabot[bot] in #951
  • [CI]: Bump astral-sh/setup-uv from 6 to 7 by @dependabot[bot] in #952
  • [CI]: Bump actions/github-script from 7 to 8 by @dependabot[bot] in #954
  • [CI]: Bump actions/checkout from 2 to 5 by @dependabot[bot] in #953
  • [TileOp] Implement WGMMA for T.gemm_v2 by @LeiWang1999 in #813
  • [Docs] add CODE_OF_CONDUCT.md by @XuehaiPan in #965
  • [Example] Add support for bfloat16 and user-defined sm_scale in attention sink examples by @Rachmanino in #924
  • [Bugfix] Do not force inline let stmt by @LeiWang1999 in #947
  • [CI] add pre-commit integration by @XuehaiPan in #955
  • [Doc] Install docs add docker install method by @BBuf in #961
  • [Bugfix] Fix dummy kernel compliation by @SiriusNEO in #962
  • [CI][Refactor] Refactor non-test CI workflow files by @XuehaiPan in #971
  • [TileOp] Implememt CumSum1D by @LeiWang1999 in #978
  • [Language] Enhance T.alloc_var for AugAssign and AnnAsign by @LeiWang1999 in #979
  • [Refactor] Refactor Pass InjectFenceProxy and expose some warp group primitives in frontend by @LeiWang1999 in #977
  • [Typo] Remove debug print by @LeiWang1999 in #980
  • [Bugfix] Use access_ptr("r") instead of access_ptr("w") for correct pipeline analysis by @LeiWang1999 in #983
  • [Feature][Example] Support TMA reduce operation and update GQA bwd example by @chengyupku in #969
  • [Bugfix] Add NVIDIA HPC SDK support in CUDA detection (#974) by @Degeneracy-Evil in #976
  • [BugFix] Robust gemm policy for sparse_mla_fwd in Hopper and Ada Lovelace architectures by @tzj-fxz in #984
  • [Bugfix] Fallback torch.accelerator.synchronize() to torch.cuda.synchronize() by @yyttt6 in #987
  • [Bugfix]:Fix atomicadd auto vectorize identify var error by @yyttt6 in #883
  • [CI] Speed up sparse tensor core test via vectorized generating sparse data by @LeiWang1999 in #1009
  • [Build] Migrate to scikit-build-core by @oraluben in #939
  • [CI] Removes redundant environment variable by @Cunxiao2002 in #1020
  • [Transform] Migrate LowerIntrin from tvm into tilelang by @LeiWang1999 in #999
  • [Lint] Prefer American English spelling by @XuehaiPan in #1022
  • [Build] Prefer libs from local build dir by @oraluben in #1027
  • [Language] Support Consequential assignments like 'a = b = c = 1' by @LeiWang1999 in #992
  • [CI] Removes debug print statements from the example. by @Cunxiao2002 in #1030
  • [Enhancement] Update abs function for half_t and bfloat_t to use cutlass implementation by @Rachmanino in #1023
  • [Bugfix] Recover code for flexible parallel by @LeiWang1999 in #1032
  • [CI] Disable buggy(maybe) warp specialized kernel ci test for H20 by @LeiWang1999 in #1033
  • [TIR] Revert some changes of Pass LowerIntrin by @LeiWang1999 in #1035
  • [Env] Optimize the mechanism for locating TL_LIBS by @LeiWang1999 in #1038
  • [CUDA] Add pack functions for FP8 types by @LJC00118 in #967
  • [Language] Expose T.get_warp_idx_sync and T.shuffle_elect for efficient thread election by @LeiWang1999 in #989
  • [AMD] fix bug&add amd fp8 examples by @Alex4210987 in #966
  • [CI][Refactor] Merge test CI workflow files into one by @XuehaiPan in #973
  • [BugFix] Phaseout dependency of Triton in sink examples to make CI happy by @Rachmanino in #1045
  • [Refactor] Use has_simt_copy to decide whether to insert set_max_nreg by @chengyupku in #982
  • [Feature]: Add test for atomicadd auto vectorize and remove useless code by @yyttt6 in #1019
  • Allow mma gemm for all cuda arch by @oraluben in #1047
  • [Bugfix] Improves compatibility when checking for MPS availability in different PyTorch builds. by @LeiWang1999 in #1051
  • [CI] Fix ROCm CI by @XuehaiPan in #1043
  • [Enhancement] Add support for symbolic dimensions in Cython kernel adapter and improve static shape validation in wrapper by @Rachmanino in #1024
  • Automatically initialize submodule if missing by @LeiWang1999 in #1052
  • [Enhancement] Remove constraint requiring last dimension stride to be 1 by @LJC00118 in #1040
  • [CI] Disable autofix for pre-commit CI by @LeiWang1999 in #1053
  • [Enhancement] Improve CUDA compiler detection in CMake by @LJC00118 in #1054
  • [Enhancement] Introduce a workaround for layout inference for local buffer store by @LeiWang1999 in #1055
  • [Refactor] Refactor Pass LegalizeSafeMemoryAccess to support recursive load/store rewrite by @SiriusNEO in #1050
  • Making version parser more robust against missing or unavailable metadata by @LeiWang1999 in #1061
  • [DOC] Add document for develop with PYTHONPATH by @LeiWang1999 in #1062
  • [CI]:Reduce test shapes to avoid OOM errors during CI. by @yyttt6 in #1060
  • [Benchmark] Add H800 SXM Benchmark results by @LeiWang1999 in #1063
  • [Misc] Add GitHub issue templates by @XuehaiPan in #1057
  • [Refactor][Example] Update linear attention examples and add tests by @Rachmanino in #1010
  • [Enhancement] Deprecate split&sum in attn bwd examples on Hopper by @Rachmanino in #1065
  • [Benchmark] Add matmul FP16 benchmark results by @LeiWang1999 in #1067
  • [CI]: Bump actions/checkout from 4 to 5 by @dependabot[bot] in #1070
  • [Example] Update GQA varlen fwd and MHA varlen fwd by @chengyupku in #1071
  • [Parallel] Support T.Parallel with dynamic extents by @LeiWang1999 in #990
  • [Layout] Utilizing IsEqual instead of StructuralEqual by @LeiWang1999 in #1073
  • [Cache] raise errors for tileang.clear_cache() by @LeiWang1999 in #1077
  • [Feature] Support Reduce operators for bitwise and/or/xor by @tzj-fxz in #1074
  • [Autotune] Add autotune coverage for symbolic M and normalize cache key by @LeiWang1999 in #1075
  • [Language] Recommend using T.dynamic instead of T.symbolic by @LeiWang1999 in #1076
  • [Language] Efficient T.reduce_ with shared memory input/output by @LeiWang1999 in #1080
  • [Bugfix] Fix missing reg alloc in custom warp specialization by @chengyupku in #1084
  • [Enhancement] Update async intrinsic handling in inject_fence_proxy by @Rachmanino in #1068
  • [Feature] Add GQA backward kernel with varlen input by @tzj-fxz in #1082
  • [BugFix] Add memory order argument for non-vectorized atomic add by @tzj-fxz in #1081
  • [Refactor] Rename cython output to tilelang_cython and relocate its path by @LeiWang1999 in #1086
  • [Target] Enhance target selection helpers and documentation by @LeiWang1999 in #1085
  • [Cleanup] Remove tilelang.disable_cache() calls from examples and tests by @Rachmanino in #1088
  • [PassConfig] Introduce PassConfig TL_STORAGE_REWRITE_DETECT_INPLACE by @LeiWang1999 in #1089
  • [Language] Support tilelang alloc_var(dtype, init=x) by @LeiWang1999 in #1092
  • [Bugfix] Fix missing host cuTensorMapEncodeIm2col call by @chengyupku in #1094
  • [GQA] Add regional atomic add to slightly boost performance by @tzj-fxz in #1093
  • [Example] Add block level high performance gemv example by @LeiWang1999 in #1097
  • [Refactor] Optimize debug message for parallel inference by @LeiWang1999 in #1096
  • [CI][Lint] Retire format.sh and add clang-tidy to GHA workflow by @XuehaiPan in #1044
  • [Refactor] Use forceinline in ldmatrix and update mamba scan kernel by @chengyupku in #1104
  • [Maint] Update uncommitted change detection command in format.sh by @XuehaiPan in #1102
  • [Benchmark] Add Mamba2_chunk_scan benchmark by @chengyupku in #1109
  • [Benchmark] Update Mamba2_chunk_scan benchmark by @chengyupku in #1110
  • [Lint] Enable pyupgrade linter in ruff by @oraluben in #963
  • [Refactor] Improve scalar handling in CopyNode and update loop partition dtype logi by @LeiWang1999 in #1111
  • [Feature] Enhance vectorized conversion support in CUDA codegen by @Rachmanino in #1095
  • [Feature] Support None type as input for T.ptr and T.Tensor by @xwhzz in #1114
  • [Bugfix] Resolve mixed stride dtype issue (inconsistent int32/int64 values) by @LeiWang1999 in #1119
  • [Feature] Add memory_order PTX for vectorized atomic add by @tzj-fxz in #1112
  • [CI]: Bump actions/upload-artifact from 4 to 5 by @dependabot[bot] in #1128
  • [CI]: Bump actions/download-artifact from 5 to 6 by @dependabot[bot] in #1127
  • [Enhancement] Add missing fence_barrier_init primitive after mbarrier init by @chengyupku in #1121
  • [Feature]:Add device assert by @yyttt6 in #1116
  • [Build][CI] Build and test SDist in release CI by @XuehaiPan in #1098
  • [Benchmark] Update triton and helion baselines in mamba-chuk-scan by @chengyupku in #1131
  • Add int2 and longlong4 pack functions by @LJC00118 in #1129
  • [BugFix] Add memory order and testing script for split version GQA bwd kernel by @tzj-fxz in #1100
  • [Bugfix] Correctly construct the argument list for atomic add based on the vector size by @LeiWang1999 in #1137
  • [AMD] Supoort T.gemm_v2 for AMD Backend by @Paran0idy in #1136
  • [BugFix] alloc_var init failed to handle complex expression by @kurisu6912 in #1144
  • [Refactor] Remove amd gemm_v2 tests by @LeiWang1999 in #1149
  • [BugFix] Implement bfloat16 support in CUDA code generation with min/max functions and inf/nan values by @Rachmanino in #1143
  • [Bugfix] Implement classic arena algorithm for shmem merge and WAW conflict detection by @LeiWang1999 in #1146
  • [CI] allow dirty workspace for format.sh and introduce loop carry thread sync unit test by @LeiWang1999 in #1153
  • [CI] use Python urllib to download file instead of Wget by @XuehaiPan in #1154
  • [BugFix] Correct direct copy from bf16 to fp8 by @Cunxiao2002 in #1090
  • [Refactor]:Move device_assert from extern_call to intrin_call by @yyttt6 in #1134
  • [Enhancement] Enhance Cast operations Vectorization by @LJC00118 in #1156
  • [Bugfix] Enhance LetStmt handling in Vectorize Loop Pass by @LeiWang1999 in #1159
  • [Release] Bump version to v0.1.6.post2 by @LeiWang1999 in #1160

New Contributors

Full Changelog: v0.1.6.post1...v0.1.6.post2