You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Commit da2c1a2
Browse filesBrowse the repository at this point in the historyBrowse files
vectorize cooperative_tensor load/store to use 4-wide memory ops
Use half4/float4 pointer casts for cooperative_tensor element loads
and stores instead of per-element scalar loops. Each 16x16 fragment
load now emits two 4-wide reads (2 rows × 4 cols) per fragment
instead of an 8-iteration scalar loop.
0 commit comments