Repository navigation
v0.1.6.post2
·
1206 commits
to main
since this release
The Last Release for Python 3.8 (without tvm-ffi) 馃殌
What's Changed
- [Analyzer] Enhance ConstIntBoundAnalyzer and IntervalSet with modular set analysis by @LeiWang1999 in #856
- [Doc] Optimize the quickstart guide for clarity and not just for CUDA by @LeiWang1999 in #858
- [TMA] Bugfix when a shared buffer is both issued with tma store and tma load by @LeiWang1999 in #857
- [AMD][MLA] Fix mla autotune for rocm by @LeiWang1999 in #861
- [Bugfix] Ensure correct handling for cases where
seq_q<seq_kvin flash attention examples by @Rachmanino in #864 - [AMD] refactor MatrixCoreIntrinEmitter by @Paran0idy in #860
- [Feat] Add fast sine and cosine definitions in CUDA templates by @Rachmanino in #865
- [Layout] Support layout forward with multi dimension by @LeiWang1999 in #867
- [Autotune][Conv] optimize convolution examples to use autotune by @LeiWang1999 in #866
- [Example] Add examples to support efficient attention sink forward process by @Rachmanino in #853
- [Parser] Adapt Parser to work with Python 3.8 in some cases by @LeiWang1999 in #869
- [Fix] Fix bug 0905: tilelang doesn't vectorize
B[i,j] = c[i] + A[i,j]by @kurisu6912 in #798 - [Language] Support sequence comparisons by @LeiWang1999 in #872
- [Language] Support loop_break primitive by @chengyupku in #873
- [Bugfix] Use
ExprDeepEqualinstead ofStructuralEqualwhen merge consecutive If stmt by @LeiWang1999 in #876 - [Language] Support atomic add with ret by @LeiWang1999 in #870
- [Cython] Remove an incorrect check by @LJC00118 in #880
- Update amd_ci.yml by @Alex4210987 in #881
- [FastMath] Disable default TVM fastmath intrinsic dispatch and add explicit fastmath op to invoke by @LeiWang1999 in #875
- [Example] Add efficient attention sink backward implementations and tests by @Rachmanino in #877
- [Precision] Introduce
T.ieee_rsqrtand related high precision op by @LeiWang1999 in #882 - [Dist] Provide an option to include commit ID in version by @LeiWang1999 in #884
- [Example] Optimize sink attention forward via swizzled layout and report benchmark results by @Rachmanino in #885
- [Layout] Introduce Flexible Parallel to Support T.serial and local buffers inside T.Parallel loop by @LeiWang1999 in #844
- [Bugfix][Enhancement] Fix a bug in previous commit and enhance cuda backend by @Hamerlate in #887
- [Bugfix] Fix CopyNode Lower method to include disable_tma flag in GetCopyInst by @Rachmanino in #888
- [Layout] Fix plot layout by @Paran0idy in #890
- [Example] Add example by @LeiWang1999 in #894
- [News] Add announcement of support for Huawei Ascend chips by @xwhzz in #895
- [Example] Add sparse mla examples by @LeiWang1999 in #896
- [Typo] Fix backend name for Huawei Ascend by @xwhzz in #898
- [CI] Legalize math related test by @LeiWang1999 in #899
- [Bugfix] Fix flops comp and softmax scale in mla by @Edenzzzz in #900
- [Example] Specify a fixed commit for the flash-linear-attention repository and optimize nsa examples by @LeiWang1999 in #913
- [CI] optimize CI time for sparse gemm by @botbw in #906
- [Enhancement] Include compile flags into the hash key of cached kernels by @Rachmanino in #911
- [Bugfix] Fix saving kernel source code where JITKernel.artifact is None by @zjudmd1015 in #921
- [CI] Refactor import paths in dequantization examples to use dequantize_utils by @LeiWang1999 in #914
- [Example] Add MLA decode ws example by @chengyupku in #928
- [CI] Fix documentation runner by adding 'nvidia' tag by @xwhzz in #927
- [Layout] Strict annotate completed replicated layout for fragment with constant index by @LeiWang1999 in #929
- [Bugfix] Fix tensor memory copy layout by @Hamerlate in #933
- [Example] Optimize online_softmax example by @lijinpei in #934
- [Example] Add correctness assert into dsa example by @LeiWang1999 in #937
- [Enhancement] Enhance and add new GQA backward examples for Hopper by @Rachmanino in #930
- [Enhancement] Fix lint to improve grouped GEMM performance with TMA by @Cunxiao2002 in #938
- [Example] Introduce split+sum template, and optimize
atomic_addperformance for bwd examples by @LeiWang1999 in #940 - [Example] Disable TMA and enable FastMath for NSA Examples (#941) by @LeiWang1999 in #941
- [Example] Revert the atomic/split&sum templates in MHA backward examples by @Rachmanino in #943
- [Example] Add sparse mla bwd example for deepseek_v32 by @Zhichenzzz in #919
- [Profiler]Adds CUPTI profiler support by @Cunxiao2002 in #936
- [Enhancement] Support Copy for Buffer Load witih scalar indices by @LeiWang1999 in #946
- [Code Style] Refine nvrtc compile related check style by @BBuf in #945
- [Backend] Add metal backend by @oraluben in #799
- [CI] enable dependabot for GHA workflows by @XuehaiPan in #950
- Modify the SM architecture number to support Thor鈥檚 sm110. by @iloveai8086 in #957
- [CI] auto-cancel in-progress PR CI when new commits are pushed by @XuehaiPan in #956
- [bug] fix type object is not subscriptable in py38 by @BBuf in #959
- [Bugfix][Doc] Add astroid version constraint to requirements.txt by @xwhzz in #958
- [CI]: Bump actions/setup-python from 2 to 6 by @dependabot[bot] in #951
- [CI]: Bump astral-sh/setup-uv from 6 to 7 by @dependabot[bot] in #952
- [CI]: Bump actions/github-script from 7 to 8 by @dependabot[bot] in #954
- [CI]: Bump actions/checkout from 2 to 5 by @dependabot[bot] in #953
- [TileOp] Implement WGMMA for T.gemm_v2 by @LeiWang1999 in #813
- [Docs] add CODE_OF_CONDUCT.md by @XuehaiPan in #965
- [Example] Add support for
bfloat16and user-definedsm_scalein attention sink examples by @Rachmanino in #924 - [Bugfix] Do not force inline let stmt by @LeiWang1999 in #947
- [CI] add
pre-commitintegration by @XuehaiPan in #955 - [Doc] Install docs add docker install method by @BBuf in #961
- [Bugfix] Fix dummy kernel compliation by @SiriusNEO in #962
- [CI][Refactor] Refactor non-test CI workflow files by @XuehaiPan in #971
- [TileOp] Implememt
CumSum1Dby @LeiWang1999 in #978 - [Language] Enhance
T.alloc_varfor AugAssign and AnnAsign by @LeiWang1999 in #979 - [Refactor] Refactor Pass
InjectFenceProxyand expose some warp group primitives in frontend by @LeiWang1999 in #977 - [Typo] Remove debug print by @LeiWang1999 in #980
- [Bugfix] Use
access_ptr("r")instead ofaccess_ptr("w")for correct pipeline analysis by @LeiWang1999 in #983 - [Feature][Example] Support TMA reduce operation and update GQA bwd example by @chengyupku in #969
- [Bugfix] Add NVIDIA HPC SDK support in CUDA detection (#974) by @Degeneracy-Evil in #976
- [BugFix] Robust gemm policy for sparse_mla_fwd in Hopper and Ada Lovelace architectures by @tzj-fxz in #984
- [Bugfix] Fallback
torch.accelerator.synchronize()totorch.cuda.synchronize()by @yyttt6 in #987 - [Bugfix]:Fix atomicadd auto vectorize identify var error by @yyttt6 in #883
- [CI] Speed up sparse tensor core test via vectorized generating sparse data by @LeiWang1999 in #1009
- [Build] Migrate to scikit-build-core by @oraluben in #939
- [CI] Removes redundant environment variable by @Cunxiao2002 in #1020
- [Transform] Migrate
LowerIntrinfrom tvm into tilelang by @LeiWang1999 in #999 - [Lint] Prefer American English spelling by @XuehaiPan in #1022
- [Build] Prefer libs from local build dir by @oraluben in #1027
- [Language] Support Consequential assignments like 'a = b = c = 1' by @LeiWang1999 in #992
- [CI] Removes debug print statements from the example. by @Cunxiao2002 in #1030
- [Enhancement] Update abs function for half_t and bfloat_t to use cutlass implementation by @Rachmanino in #1023
- [Bugfix] Recover code for flexible parallel by @LeiWang1999 in #1032
- [CI] Disable buggy(maybe) warp specialized kernel ci test for H20 by @LeiWang1999 in #1033
- [TIR] Revert some changes of Pass
LowerIntrinby @LeiWang1999 in #1035 - [Env] Optimize the mechanism for locating
TL_LIBSby @LeiWang1999 in #1038 - [CUDA] Add pack functions for FP8 types by @LJC00118 in #967
- [Language] Expose
T.get_warp_idx_syncandT.shuffle_electfor efficient thread election by @LeiWang1999 in #989 - [AMD] fix bug&add amd fp8 examples by @Alex4210987 in #966
- [CI][Refactor] Merge test CI workflow files into one by @XuehaiPan in #973
- [BugFix] Phaseout dependency of Triton in sink examples to make CI happy by @Rachmanino in #1045
- [Refactor] Use
has_simt_copyto decide whether to insertset_max_nregby @chengyupku in #982 - [Feature]: Add test for atomicadd auto vectorize and remove useless code by @yyttt6 in #1019
- Allow mma gemm for all cuda arch by @oraluben in #1047
- [Bugfix] Improves compatibility when checking for MPS availability in different PyTorch builds. by @LeiWang1999 in #1051
- [CI] Fix ROCm CI by @XuehaiPan in #1043
- [Enhancement] Add support for symbolic dimensions in Cython kernel adapter and improve static shape validation in wrapper by @Rachmanino in #1024
- Automatically initialize submodule if missing by @LeiWang1999 in #1052
- [Enhancement] Remove constraint requiring last dimension stride to be 1 by @LJC00118 in #1040
- [CI] Disable autofix for pre-commit CI by @LeiWang1999 in #1053
- [Enhancement] Improve CUDA compiler detection in CMake by @LJC00118 in #1054
- [Enhancement] Introduce a workaround for layout inference for local buffer store by @LeiWang1999 in #1055
- [Refactor] Refactor Pass
LegalizeSafeMemoryAccessto support recursive load/store rewrite by @SiriusNEO in #1050 - Making version parser more robust against missing or unavailable metadata by @LeiWang1999 in #1061
- [DOC] Add document for develop with PYTHONPATH by @LeiWang1999 in #1062
- [CI]:Reduce test shapes to avoid OOM errors during CI. by @yyttt6 in #1060
- [Benchmark] Add H800 SXM Benchmark results by @LeiWang1999 in #1063
- [Misc] Add GitHub issue templates by @XuehaiPan in #1057
- [Refactor][Example] Update linear attention examples and add tests by @Rachmanino in #1010
- [Enhancement] Deprecate split&sum in attn bwd examples on Hopper by @Rachmanino in #1065
- [Benchmark] Add matmul FP16 benchmark results by @LeiWang1999 in #1067
- [CI]: Bump actions/checkout from 4 to 5 by @dependabot[bot] in #1070
- [Example] Update GQA varlen fwd and MHA varlen fwd by @chengyupku in #1071
- [Parallel] Support
T.Parallelwith dynamic extents by @LeiWang1999 in #990 - [Layout] Utilizing IsEqual instead of StructuralEqual by @LeiWang1999 in #1073
- [Cache] raise errors for
tileang.clear_cache()by @LeiWang1999 in #1077 - [Feature] Support Reduce operators for bitwise and/or/xor by @tzj-fxz in #1074
- [Autotune] Add autotune coverage for symbolic M and normalize cache key by @LeiWang1999 in #1075
- [Language] Recommend using
T.dynamicinstead ofT.symbolicby @LeiWang1999 in #1076 - [Language] Efficient
T.reduce_with shared memory input/output by @LeiWang1999 in #1080 - [Bugfix] Fix missing reg alloc in custom warp specialization by @chengyupku in #1084
- [Enhancement] Update async intrinsic handling in inject_fence_proxy by @Rachmanino in #1068
- [Feature] Add GQA backward kernel with varlen input by @tzj-fxz in #1082
- [BugFix] Add memory order argument for non-vectorized atomic add by @tzj-fxz in #1081
- [Refactor] Rename cython output to
tilelang_cythonand relocate its path by @LeiWang1999 in #1086 - [Target] Enhance target selection helpers and documentation by @LeiWang1999 in #1085
- [Cleanup] Remove
tilelang.disable_cache()calls from examples and tests by @Rachmanino in #1088 - [PassConfig] Introduce PassConfig
TL_STORAGE_REWRITE_DETECT_INPLACEby @LeiWang1999 in #1089 - [Language] Support tilelang
alloc_var(dtype, init=x)by @LeiWang1999 in #1092 - [Bugfix] Fix missing host
cuTensorMapEncodeIm2colcall by @chengyupku in #1094 - [GQA] Add regional atomic add to slightly boost performance by @tzj-fxz in #1093
- [Example] Add block level high performance gemv example by @LeiWang1999 in #1097
- [Refactor] Optimize debug message for parallel inference by @LeiWang1999 in #1096
- [CI][Lint] Retire
format.shand addclang-tidyto GHA workflow by @XuehaiPan in #1044 - [Refactor] Use forceinline in
ldmatrixand update mamba scan kernel by @chengyupku in #1104 - [Maint] Update uncommitted change detection command in
format.shby @XuehaiPan in #1102 - [Benchmark] Add Mamba2_chunk_scan benchmark by @chengyupku in #1109
- [Benchmark] Update Mamba2_chunk_scan benchmark by @chengyupku in #1110
- [Lint] Enable pyupgrade linter in ruff by @oraluben in #963
- [Refactor] Improve scalar handling in CopyNode and update loop partition dtype logi by @LeiWang1999 in #1111
- [Feature] Enhance vectorized conversion support in CUDA codegen by @Rachmanino in #1095
- [Feature] Support None type as input for
T.ptrandT.Tensorby @xwhzz in #1114 - [Bugfix] Resolve mixed stride dtype issue (inconsistent int32/int64 values) by @LeiWang1999 in #1119
- [Feature] Add memory_order PTX for vectorized atomic add by @tzj-fxz in #1112
- [CI]: Bump actions/upload-artifact from 4 to 5 by @dependabot[bot] in #1128
- [CI]: Bump actions/download-artifact from 5 to 6 by @dependabot[bot] in #1127
- [Enhancement] Add missing
fence_barrier_initprimitive after mbarrier init by @chengyupku in #1121 - [Feature]:Add device assert by @yyttt6 in #1116
- [Build][CI] Build and test SDist in release CI by @XuehaiPan in #1098
- [Benchmark] Update triton and helion baselines in mamba-chuk-scan by @chengyupku in #1131
- Add int2 and longlong4 pack functions by @LJC00118 in #1129
- [BugFix] Add memory order and testing script for split version GQA bwd kernel by @tzj-fxz in #1100
- [Bugfix] Correctly construct the argument list for atomic add based on the vector size by @LeiWang1999 in #1137
- [AMD] Supoort T.gemm_v2 for AMD Backend by @Paran0idy in #1136
- [BugFix] alloc_var init failed to handle complex expression by @kurisu6912 in #1144
- [Refactor] Remove amd gemm_v2 tests by @LeiWang1999 in #1149
- [BugFix] Implement bfloat16 support in CUDA code generation with min/max functions and inf/nan values by @Rachmanino in #1143
- [Bugfix] Implement classic arena algorithm for shmem merge and WAW conflict detection by @LeiWang1999 in #1146
- [CI] allow dirty workspace for
format.shand introduce loop carry thread sync unit test by @LeiWang1999 in #1153 - [CI] use Python urllib to download file instead of Wget by @XuehaiPan in #1154
- [BugFix] Correct direct copy from bf16 to fp8 by @Cunxiao2002 in #1090
- [Refactor]:Move device_assert from extern_call to intrin_call by @yyttt6 in #1134
- [Enhancement] Enhance Cast operations Vectorization by @LJC00118 in #1156
- [Bugfix] Enhance LetStmt handling in Vectorize Loop Pass by @LeiWang1999 in #1159
- [Release] Bump version to v0.1.6.post2 by @LeiWang1999 in #1160
New Contributors
- @LJC00118 made their first contribution in #880
- @Edenzzzz made their first contribution in #900
- @zjudmd1015 made their first contribution in #921
- @lijinpei made their first contribution in #934
- @Zhichenzzz made their first contribution in #919
- @BBuf made their first contribution in #945
- @XuehaiPan made their first contribution in #950
- @iloveai8086 made their first contribution in #957
- @Degeneracy-Evil made their first contribution in #976
Full Changelog: v0.1.6.post1...v0.1.6.post2