Repository navigation
Add Cuda Vector Types - #7949
Conversation
|
gpuci run tests |
gmarkall
left a comment
There was a problem hiding this comment.
I've partially reviewed this so far, but I think it's worth posting some comments now with more to follow - there are some comments / suggestions on the diff.
Can we support casting between vectors that are essentially the same? For example, the following:
from numba import cuda, types
def f(c):
if c:
return cuda.int2(2, 2)
else:
return cuda.long2(2, 2)
cuda.compile_ptx(f, (types.boolean,), device=True)presently gives:
numba.core.errors.TypingError: Failed in cuda mode pipeline (step: nopython frontend)
Can't unify return type from the following types: int2, long2
Return of: IR name '$16return_value.5', type 'int2', location:
File "cast_same.py", line 6:
def f(c):
<source elided>
if c:
return cuda.int2(2, 2)
^
Return of: IR name '$28return_value.5', type 'long2', location:
File "cast_same.py", line 8:
def f(c):
<source elided>
else:
return cuda.long2(2, 2)
^
Do we need the once decorator in vector_types.py? Is it to prevent a circular import? If it is needed, can we just call the function initialize()? (I feel like the "once" part of it is an implementation detail, in a sense).
I've still yet to properly look at the majority of vector_types.py and the tests - I'll look at these in the next pass.
Co-authored-by: Graham Markall <535640+gmarkall@users.noreply.github.com>
…fea/cuda_vector_types
To clarify -
|
@gmarkall and I discussed offline on this and there are two solutions: one is to implement the same integer types as aliases and the other is to implement the casting rules. We favored implemented as aliases; it aligns more with how numba types are implemented. The main cons of this solutions are that the kernels won't be portable across platforms. For example To address this concern - we plan to define the vector type for python kernels with a new set of names more aligned with numpy's naming conventions: it would look like In addition, all names from native CUDA will be aliased on this new naming convention (based on machine definition of |
Co-authored-by: Graham Markall <535640+gmarkall@users.noreply.github.com>
…lated vector types
|
gpuci run tests |
gmarkall
left a comment
There was a problem hiding this comment.
Many thanks for the updates - just some very small comments / questions on the diff.
|
gpuci run tests |
gmarkall
left a comment
There was a problem hiding this comment.
Many thanks for addressing the comments!
|
@esc / @stuartarchibald Could this have a CUDA buildfarm smoketest please? |
|
Thanks @esc, this passed. |
|
Thanks for the buildfarm run @esc / @stuartarchibald. From OOB discussion, there was some concern about the length of time the Windows builders in the buildfarm took - however, I tested this on my machine locally (Win11, RTX 6000, Xeon 6128, 96GB RAM) and the increase in test time on Windows was negligible on a modern quiescent machine - it is likely that what was observed was just the general slowness of the Windows buildfarm machines. |
This is the initial introduction of cuda vector types. This PR adds support to vector type creation and readouts. Note that not only vector types can be constructed a number of primitive types, but also with all valid combinations of vector type and primitives. Such as:
Contributes to #7943 .