Repository navigation
[Metal] Add Metal GEMM support with simdgroup_matrix MMA #1869
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Merged
Merged
Changes from 1 commit
Commits
Show all changes
16 commits
Select commit
Hold shift + click to select a range
4b015bf
[Metal] Add Metal GEMM support with simdgroup_matrix MMA
oraluben 9e45e47
Merge branch 'main' of https://github.com/tile-ai/tilelang into metal…
LeiWang1999 a3fdb11
Move Metal buffer helpers to Metal backend
LeiWang1999 950d009
fix(metal): remove layout_inference bypass, fix build & clarify docs
oraluben ce6e4c2
Merge upstream/main - fix backend import path for metal
oraluben eaa9569
Merge main into metal-gemm
oraluben 0eb2865
fix: tir->tirx migration build fixes for metal backend
oraluben ab88cae
fix: improve metal_fragment_to_simdgroup buffer remapping
oraluben 8e8e83b
lint
oraluben 5bff0ee
fix: add full-body var subst in MetalFragmentToSimdgroup to reach gem…
oraluben cd48d6a
fix: add missing <cmath> include for std::isinf/isnan in codegen_metal
oraluben 4707ede
Merge branch 'main' into metal-gemm
oraluben 7f7adfd
[Metal] Fix review findings: dedup warp partition, kernel_only, barri…
oraluben db6c2f7
lint
oraluben 147f37c
Refactor MetalFragmentToSimdgroup import path in LowerAndLegalize fun…
LeiWang1999 b7fcad6
Remove metal_fragment_to_simdgroup.py file and update import in metal…
LeiWang1999 File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
[Metal] Fix review findings: dedup warp partition, kernel_only, barri…
…er comment, immutable gemm ops - Extract shared ComputeSquareWarpPartition to src/backend/metal/op/utils.h, used by both gemm.cc (Square policy) and copy.cc (simdgroup store lowering). - Implement kernel_only parameter in MetalKernelAdapter.get_kernel_source. - Add comment documenting that Metal's simdgroup_barrier synchronizes the full threadgroup (no per-simdgroup barrier exists). - Replace module-level mutable global in metal_fragment_to_simdgroup.py with @lru_cache and remove CUDA-only gemm ops from the set.
- Loading branch information
commit 7f7adfde43d259e4b28e288bfb4c5aef554bf6c0
There are no files selected for viewing
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
Fix the return type annotation and honour the
kernel_onlyflag.Two issues here:
Return type mismatch —
kernel_global_sourceis declaredstr | None(Line 23), so the method can returnNone, contradicting the-> strannotation. This will cause silent type errors for callers.Unused
kernel_onlyparameter — Every peer adapter branches on this flag (base.py,nvrtc/adapter.py,cython/adapter.py). Silently ignoring it here meansget_kernel_source(kernel_only=False)behaves identically tokernel_only=True, breaking the expected contract.🛠️ Proposed fix
If a non-
Noneguarantee is truly required at call sites, add an explicit assertion:🧰 Tools
🪛 Ruff (0.15.1)
[warning] 56-56: Unused method argument:
kernel_only(ARG002)
🤖 Prompt for AI Agents