Skip to content

Commit 708fd43

Browse files
[Fix][NVPTX] Cast thread index to the loop variable dtype (#20476)
NVPTX GetThreadIndex returned the raw i32 from the tid/ctaid intrinsics even when the bound loop var is int64. Kernels with symbolic int64 extents then emitted mixed-width LLVM IR and failed module verification. This casts the result to the loop var dtype, matching the AMDGPU backend.
1 parent 076e83f commit 708fd43

1 file changed

Lines changed: 2 additions & 1 deletion

File tree

‎src/backend/cuda/codegen/llvm/codegen_nvptx.cc‎

Lines changed: 2 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -169,7 +169,8 @@ class CodeGenNVPTX : public CodeGenLLVM {
169169
#else
170170
llvm::Function* f = llvm::Intrinsic::getDeclaration(module_.get(), intrin_id);
171171
#endif
172-
return builder_->CreateCall(f, {});
172+
llvm::Value* result = builder_->CreateCall(f, {});
173+
return this->CreateCast(PrimType::Int(32), iv->var.ty(), result);
173174
}
174175

175176
llvm::Value* CreateStorageSync(const CallNode* op) final {

0 commit comments

Comments
 (0)