Skip to content

Add FP16 support for CUDA - #7460

Merged
sklam merged 35 commits into
numba:masterfrom
testhound:testhound/fp16-support
Dec 15, 2021
Merged

sklam merged 35 commits into
numba:masterfrom
testhound:testhound/fp16-support

Conversation

@testhound

Copy link
Copy Markdown
Contributor

This PR begins the process of adding 16-bit floating point support for the cuda target. This PR adds support for the following 16-bit fp operators: add, subtract, multiply, multiply accumulate, negate and absolute value. The divide operator was omitted as it is more involved and will be included in the next PR.

@sklam

sklam commented Oct 6, 2021 •

Copy link
Copy Markdown
Member

Looks like the CI failure is because the fp16 intrinsics is not available in cudasim

@sklam sklam added 2 - In Progress CUDA CUDA related issue/PR labels Oct 6, 2021
@testhound

Copy link
Copy Markdown
Contributor Author

Looks like the CI failure is because the fp16 intrinsics is not available in cudasim

@sklam pushed a patch for cudasim.

@gmarkall gmarkall added the Effort - medium Medium size effort needed label Oct 6, 2021
@gmarkall gmarkall added this to the Numba 0.55 RC milestone Oct 6, 2021
@gmarkall

gmarkall commented Oct 7, 2021

Copy link
Copy Markdown
Member

The following:

from numba import cuda
import numpy as np


@cuda.jit
def fadd(r, x, y):
    r[0] = cuda.fp16.hadd(x[0], y[0])

x = np.ones(1, dtype=np.float16) * 2.5
y = np.ones(1, dtype=np.float16) * 3.7
r = np.zeros_like(x)

fadd[1, 1](r, x, y)

yields

Traceback (most recent call last):
  File "/home/gmarkall/numbadev/numba/numba/core/typing/typeof.py", line 237, in _typeof_ndarray
    dtype = numpy_support.from_dtype(val.dtype)
  File "/home/gmarkall/numbadev/numba/numba/np/numpy_support.py", line 113, in from_dtype
    raise NotImplementedError(dtype)
NotImplementedError: float16

During handling of the above exception, another exception occurred:

Traceback (most recent call last):
  File "/home/gmarkall/numbadev/issues/7460/repro.py", line 13, in <module>
    fadd[1, 1](r, x, y)
  File "/home/gmarkall/numbadev/numba/numba/cuda/compiler.py", line 878, in __call__
    return self.dispatcher.call(args, self.griddim, self.blockdim,
  File "/home/gmarkall/numbadev/numba/numba/cuda/compiler.py", line 1011, in call
    kernel = _dispatcher.Dispatcher._cuda_call(self, *args)
  File "/home/gmarkall/numbadev/numba/numba/cuda/compiler.py", line 1036, in typeof_pyval
    return typeof(val, Purpose.argument)
  File "/home/gmarkall/numbadev/numba/numba/core/typing/typeof.py", line 31, in typeof
    ty = typeof_impl(val, c)
  File "/home/gmarkall/miniconda3/envs/numbanp120/lib/python3.9/functools.py", line 877, in wrapper
    return dispatch(args[0].__class__)(*args, **kw)
  File "/home/gmarkall/numbadev/numba/numba/core/typing/typeof.py", line 239, in _typeof_ndarray
    raise ValueError("Unsupported array dtype: %s" % (val.dtype,))
ValueError: Unsupported array dtype: float16

Am I doing something we don't expect to support yet here?

@testhound

Copy link
Copy Markdown
Contributor Author

The following:

from numba import cuda
import numpy as np


@cuda.jit
def fadd(r, x, y):
    r[0] = cuda.fp16.hadd(x[0], y[0])

x = np.ones(1, dtype=np.float16) * 2.5
y = np.ones(1, dtype=np.float16) * 3.7
r = np.zeros_like(x)

fadd[1, 1](r, x, y)

yields

Traceback (most recent call last):
  File "/home/gmarkall/numbadev/numba/numba/core/typing/typeof.py", line 237, in _typeof_ndarray
    dtype = numpy_support.from_dtype(val.dtype)
  File "/home/gmarkall/numbadev/numba/numba/np/numpy_support.py", line 113, in from_dtype
    raise NotImplementedError(dtype)
NotImplementedError: float16

During handling of the above exception, another exception occurred:

Traceback (most recent call last):
  File "/home/gmarkall/numbadev/issues/7460/repro.py", line 13, in <module>
    fadd[1, 1](r, x, y)
  File "/home/gmarkall/numbadev/numba/numba/cuda/compiler.py", line 878, in __call__
    return self.dispatcher.call(args, self.griddim, self.blockdim,
  File "/home/gmarkall/numbadev/numba/numba/cuda/compiler.py", line 1011, in call
    kernel = _dispatcher.Dispatcher._cuda_call(self, *args)
  File "/home/gmarkall/numbadev/numba/numba/cuda/compiler.py", line 1036, in typeof_pyval
    return typeof(val, Purpose.argument)
  File "/home/gmarkall/numbadev/numba/numba/core/typing/typeof.py", line 31, in typeof
    ty = typeof_impl(val, c)
  File "/home/gmarkall/miniconda3/envs/numbanp120/lib/python3.9/functools.py", line 877, in wrapper
    return dispatch(args[0].__class__)(*args, **kw)
  File "/home/gmarkall/numbadev/numba/numba/core/typing/typeof.py", line 239, in _typeof_ndarray
    raise ValueError("Unsupported array dtype: %s" % (val.dtype,))
ValueError: Unsupported array dtype: float16

Am I doing something we don't expect to support yet here?

@gmarkall no I don't see that you are doing anything wrong. Need to dig into this.

@sklam

sklam commented Oct 8, 2021

Copy link
Copy Markdown
Member

I tracked the recent CI error to the ufunc loop selection code. The change in 95727a4#diff-32316606fefc1b740b8e2738aa2ed5fa533d7cec87b1a691bbe65ac5daab12f3R32 allow ufunc loop for float16 to be selected by there are no float16 implementation yet. The following patch stops it from selecting ufunc loop that has no existing implementation:

diff --git a/numba/np/numpy_support.py b/numba/np/numpy_support.py
index c83dc3cf9..a5530c7da 100644
--- a/numba/np/numpy_support.py
+++ b/numba/np/numpy_support.py
@@ -497,7 +497,9 @@ def ufunc_find_matching_loop(ufunc, arg_types):
                 # (e.g. float16), try other candidates
                 continue
             else:
-                return UFuncLoopSpec(inputs, outputs, candidate)
+                loopspec = UFuncLoopSpec(inputs, outputs, candidate)
+                if supported_ufunc_loop(ufunc, loopspec):
+                    return loopspec

     return None

@sklam

sklam commented Oct 12, 2021 •

Copy link
Copy Markdown
Member

Looks like I shouldn't have changed ufunc_find_matching_loop(). The failing test is expecting it to behave in a certain way. I'm going to patch it by making that function reject float16 for now---we don't have any ufunc implementation for float16 yet.

The patch:

diff --git a/numba/np/numpy_support.py b/numba/np/numpy_support.py
index a5530c7da..b8ee24135 100644
--- a/numba/np/numpy_support.py
+++ b/numba/np/numpy_support.py
@@ -455,6 +455,9 @@ def ufunc_find_matching_loop(ufunc, arg_types):
     for candidate in ufunc.types:
         ufunc_inputs = candidate[:ufunc.nin]
         ufunc_outputs = candidate[-ufunc.nout:] if ufunc.nout else []
+        if 'e' in ufunc_inputs:
+            # Skip float16 arrays since we don't have implementation for them
+            continue
         if 'O' in ufunc_inputs:
             # Skip object arrays
             continue
@@ -497,9 +500,7 @@ def ufunc_find_matching_loop(ufunc, arg_types):
                 # (e.g. float16), try other candidates
                 continue
             else:
-                loopspec = UFuncLoopSpec(inputs, outputs, candidate)
-                if supported_ufunc_loop(ufunc, loopspec):
-                    return loopspec
+                return UFuncLoopSpec(inputs, outputs, candidate)
 
     return None
 

@gmarkall gmarkall left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Many thanks for the PR! The big picture looks good - there are some comments on the diff, and a couple of other observations:

  • The changes in boxing.py, builtins.py, and cudamath.py appear to be unnecessary to support fp16 in CUDA and the currently-implemented intrinsics. The commit bfa193d undoes these changes (and a whitespace change in compiler.py).
  • The logic for lowering could be simplified and shortened - commit 484557f factors some of them and simplifies things a little (in the same way that the typing is already factored in this PR).

The above commits along with Siu's new proposed fix test OK on CI (checked in #7481) and also all passes locally for me with hardware and the simulator - feel free to add those commits to this branch if you'd like.

Comment thread docs/source/cuda-reference/kernel.rst Outdated
Comment thread docs/source/cuda-reference/kernel.rst Outdated
Comment thread docs/source/cuda-reference/kernel.rst Outdated
Comment thread numba/core/itanium_mangler.py Outdated
'unsigned long long': 'y', # unsigned __int64
'__int128': 'n',
'unsigned __int128': 'o',
'float16' : 'Dh',

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note to any other reviewers: correlates with https://itanium-cxx-abi.github.io/cxx-abi/abi.html#mangling-builtin

Comment thread numba/cuda/cudaimpl.py Outdated
Comment thread numba/cuda/cudaimpl.py Outdated
Comment thread numba/cuda/cudaimpl.py
Comment thread numba/cuda/cudamath.py Outdated
Comment thread numba/cuda/stubs.py Outdated

@gmarkall gmarkall left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Many thanks for the updates ! Things are looking good, though there are a couple of comments / questions on the diff.

The additions to the type casting rules and corresponding tests look good, and appear to be consistent with existing type casting rules. I'm not 100% sure I understand why promoting to a wider float from a narrower one is considered unsafe (e.g. the existing f32 -> f64, but I'm assuming it's part of a tradeoff somewhere to try and avoid widening things unnecessarily.

Comment thread numba/cuda/cudaimpl.py Outdated
Comment on lines +444 to +447
arg1 = args[0]
arg2 = args[1]
arg3 = args[2]
return builder.call(hfma_inline, [arg1, arg2, arg3])

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
arg1 = args[0]
arg2 = args[1]
arg3 = args[2]
return builder.call(hfma_inline, [arg1, arg2, arg3])
return builder.call(hfma_inline, args)



@skip_on_cudasim('CUDA Driver API unsupported in the simulator')
@unittest.skip

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Was this added in error?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes.

- Don't raise a lowering error if CUDA toolkit is < 10.2 when lowering
  habs - use an alternative implementation instead.
- Rename the `float16` C type to `half`.
- Swap order of some lines that could raise so the raise happens earlier
  in the function
- Simplify return type of a function declaration in
  `integer_to_float16_cast`.
@gmarkall

Copy link
Copy Markdown
Member

gpuci run tests

@gmarkall

Copy link
Copy Markdown
Member

gpuci run tests

@stuartarchibald stuartarchibald left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the updates to address the review @gmarkall, patch looks good.

@stuartarchibald stuartarchibald added 4 - Waiting on CI Review etc done, waiting for CI to finish Pending BuildFarm For PRs that have been reviewed but pending a push through our buildfarm and removed 4 - Waiting on author Waiting for author to respond to review labels Dec 14, 2021
@esc

esc commented Dec 14, 2021

Copy link
Copy Markdown
Member

Build numba_smoketest_cuda_yaml_108 has started

gmarkall
gmarkall previously approved these changes Dec 14, 2021

@gmarkall gmarkall left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM pending buildfarm.

@esc

esc commented Dec 14, 2021

Copy link
Copy Markdown
Member

Interestingly enough this breaks across all Windows builds with the following errors:

======================================================================
ERROR: test_literal_to_float16 (numba.cuda.tests.cudapy.test_casting.TestCasting) (func=<function cuda_int_literal_to_float16 at 0x000002332EB61550>)
----------------------------------------------------------------------
Traceback (most recent call last):
  File "F:\ci_envs\64\Miniconda3\envs\testenv_03f65b4d-5932-4ea8-9027-cc24ca0c5d36\lib\site-packages\numba\cuda\cudadrv\driver.py", line 2682, in add_ptx
    driver.cuLinkAddData(self.handle, enums.CU_JIT_INPUT_PTX,
  File "F:\ci_envs\64\Miniconda3\envs\testenv_03f65b4d-5932-4ea8-9027-cc24ca0c5d36\lib\site-packages\numba\cuda\cudadrv\driver.py", line 319, in safe_cuda_api_call
    self._check_ctypes_error(fname, retcode)
  File "F:\ci_envs\64\Miniconda3\envs\testenv_03f65b4d-5932-4ea8-9027-cc24ca0c5d36\lib\site-packages\numba\cuda\cudadrv\driver.py", line 384, in _check_ctypes_error
    raise CudaAPIError(retcode, msg)
numba.cuda.cudadrv.driver.CudaAPIError: [218] Call to cuLinkAddData results in UNKNOWN_CUDA_ERROR

During handling of the above exception, another exception occurred:

Traceback (most recent call last):
  File "F:\ci_envs\64\Miniconda3\envs\testenv_03f65b4d-5932-4ea8-9027-cc24ca0c5d36\lib\site-packages\numba\cuda\tests\cudapy\test_casting.py", line 172, in test_literal_to_float16
    self.assertEqual(cfunc(321), hostfunc(321))
  File "F:\ci_envs\64\Miniconda3\envs\testenv_03f65b4d-5932-4ea8-9027-cc24ca0c5d36\lib\site-packages\numba\cuda\tests\cudapy\test_casting.py", line 105, in wrapper_fn
    cuda_wrapper_fn[1, 1](argarray, resarray)
  File "F:\ci_envs\64\Miniconda3\envs\testenv_03f65b4d-5932-4ea8-9027-cc24ca0c5d36\lib\site-packages\numba\cuda\compiler.py", line 727, in __call__
    return self.dispatcher.call(args, self.griddim, self.blockdim,
  File "F:\ci_envs\64\Miniconda3\envs\testenv_03f65b4d-5932-4ea8-9027-cc24ca0c5d36\lib\site-packages\numba\cuda\compiler.py", line 915, in call
    kernel = _dispatcher.Dispatcher._cuda_call(self, *args)
  File "F:\ci_envs\64\Miniconda3\envs\testenv_03f65b4d-5932-4ea8-9027-cc24ca0c5d36\lib\site-packages\numba\cuda\compiler.py", line 923, in _compile_for_args
    return self.compile(tuple(argtypes))
  File "F:\ci_envs\64\Miniconda3\envs\testenv_03f65b4d-5932-4ea8-9027-cc24ca0c5d36\lib\site-packages\numba\cuda\compiler.py", line 1093, in compile
    kernel.bind()
  File "F:\ci_envs\64\Miniconda3\envs\testenv_03f65b4d-5932-4ea8-9027-cc24ca0c5d36\lib\site-packages\numba\cuda\compiler.py", line 466, in bind
    self._codelibrary.get_cufunc()
  File "F:\ci_envs\64\Miniconda3\envs\testenv_03f65b4d-5932-4ea8-9027-cc24ca0c5d36\lib\site-packages\numba\cuda\codegen.py", line 201, in get_cufunc
    cubin = self.get_cubin(cc=device.compute_capability)
  File "F:\ci_envs\64\Miniconda3\envs\testenv_03f65b4d-5932-4ea8-9027-cc24ca0c5d36\lib\site-packages\numba\cuda\codegen.py", line 174, in get_cubin
    linker.add_ptx(ptx.encode())
  File "F:\ci_envs\64\Miniconda3\envs\testenv_03f65b4d-5932-4ea8-9027-cc24ca0c5d36\lib\site-packages\numba\cuda\cudadrv\driver.py", line 2685, in add_ptx
    raise LinkerError("%s\n%s" % (e, self.error_log))
numba.cuda.cudadrv.driver.LinkerError: [218] Call to cuLinkAddData results in UNKNOWN_CUDA_ERROR
ptxas application ptx input, line 54; error   : Feature 'f16 arithemetic and compare instructions' requires .target sm_53 or higher
ptxas fatal   : Ptx assembly aborted due to errors

======================================================================
ERROR: test_literal_to_float16 (numba.cuda.tests.cudapy.test_casting.TestCasting) (func=<function cuda_float_literal_to_float16 at 0x000002332EB61670>)
----------------------------------------------------------------------
Traceback (most recent call last):
  File "F:\ci_envs\64\Miniconda3\envs\testenv_03f65b4d-5932-4ea8-9027-cc24ca0c5d36\lib\site-packages\numba\cuda\cudadrv\driver.py", line 2682, in add_ptx
    driver.cuLinkAddData(self.handle, enums.CU_JIT_INPUT_PTX,
  File "F:\ci_envs\64\Miniconda3\envs\testenv_03f65b4d-5932-4ea8-9027-cc24ca0c5d36\lib\site-packages\numba\cuda\cudadrv\driver.py", line 319, in safe_cuda_api_call
    self._check_ctypes_error(fname, retcode)
  File "F:\ci_envs\64\Miniconda3\envs\testenv_03f65b4d-5932-4ea8-9027-cc24ca0c5d36\lib\site-packages\numba\cuda\cudadrv\driver.py", line 384, in _check_ctypes_error
    raise CudaAPIError(retcode, msg)
numba.cuda.cudadrv.driver.CudaAPIError: [218] Call to cuLinkAddData results in UNKNOWN_CUDA_ERROR

During handling of the above exception, another exception occurred:

Traceback (most recent call last):
  File "F:\ci_envs\64\Miniconda3\envs\testenv_03f65b4d-5932-4ea8-9027-cc24ca0c5d36\lib\site-packages\numba\cuda\tests\cudapy\test_casting.py", line 172, in test_literal_to_float16
    self.assertEqual(cfunc(321), hostfunc(321))
  File "F:\ci_envs\64\Miniconda3\envs\testenv_03f65b4d-5932-4ea8-9027-cc24ca0c5d36\lib\site-packages\numba\cuda\tests\cudapy\test_casting.py", line 105, in wrapper_fn
    cuda_wrapper_fn[1, 1](argarray, resarray)
  File "F:\ci_envs\64\Miniconda3\envs\testenv_03f65b4d-5932-4ea8-9027-cc24ca0c5d36\lib\site-packages\numba\cuda\compiler.py", line 727, in __call__
    return self.dispatcher.call(args, self.griddim, self.blockdim,
  File "F:\ci_envs\64\Miniconda3\envs\testenv_03f65b4d-5932-4ea8-9027-cc24ca0c5d36\lib\site-packages\numba\cuda\compiler.py", line 915, in call
    kernel = _dispatcher.Dispatcher._cuda_call(self, *args)
  File "F:\ci_envs\64\Miniconda3\envs\testenv_03f65b4d-5932-4ea8-9027-cc24ca0c5d36\lib\site-packages\numba\cuda\compiler.py", line 923, in _compile_for_args
    return self.compile(tuple(argtypes))
  File "F:\ci_envs\64\Miniconda3\envs\testenv_03f65b4d-5932-4ea8-9027-cc24ca0c5d36\lib\site-packages\numba\cuda\compiler.py", line 1093, in compile
    kernel.bind()
  File "F:\ci_envs\64\Miniconda3\envs\testenv_03f65b4d-5932-4ea8-9027-cc24ca0c5d36\lib\site-packages\numba\cuda\compiler.py", line 466, in bind
    self._codelibrary.get_cufunc()
  File "F:\ci_envs\64\Miniconda3\envs\testenv_03f65b4d-5932-4ea8-9027-cc24ca0c5d36\lib\site-packages\numba\cuda\codegen.py", line 201, in get_cufunc
    cubin = self.get_cubin(cc=device.compute_capability)
  File "F:\ci_envs\64\Miniconda3\envs\testenv_03f65b4d-5932-4ea8-9027-cc24ca0c5d36\lib\site-packages\numba\cuda\codegen.py", line 174, in get_cubin
    linker.add_ptx(ptx.encode())
  File "F:\ci_envs\64\Miniconda3\envs\testenv_03f65b4d-5932-4ea8-9027-cc24ca0c5d36\lib\site-packages\numba\cuda\cudadrv\driver.py", line 2685, in add_ptx
    raise LinkerError("%s\n%s" % (e, self.error_log))
numba.cuda.cudadrv.driver.LinkerError: [218] Call to cuLinkAddData results in UNKNOWN_CUDA_ERROR
ptxas application ptx input, line 55; error   : Feature 'f16 arithemetic and compare instructions' requires .target sm_53 or higher
ptxas fatal   : Ptx assembly aborted due to errors

----------------------------------------------------------------------
Ran 1257 tests in 231.363s

FAILED (errors=2, skipped=53, expected failures=7)

@gmarkall
gmarkall dismissed stale reviews from stuartarchibald and themself via 20c1d4f December 14, 2021 16:40
@gmarkall

Copy link
Copy Markdown
Member

gpuci run tests

@seibert

seibert commented Dec 14, 2021

Copy link
Copy Markdown
Contributor

This shows a limitation of the internal Numba Windows test server, which currently has a sm_35 card (Tesla K40c) in it. We need to upgrade that server to a newer card anyway since NVIDIA has already dropped support for this card in the latest CUDA release. For now, I think we need to mark these tests to skip on older architectures.

@gmarkall

Copy link
Copy Markdown
Member

I've pushed a fix that skips the test if the CC is not new enough. I think we should consider dropping support for CC less than at least 5.2 after Numba 0.55.

@stuartarchibald stuartarchibald left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for addressing the issue with CC < 5.3.

@sklam

sklam commented Dec 14, 2021

Copy link
Copy Markdown
Member

BFID numba_smoketest_cuda_yaml_109

@sklam sklam added 5 - Ready to merge Review and testing done, is ready to merge BuildFarm Passed For PRs that have been through the buildfarm and passed and removed Pending BuildFarm For PRs that have been reviewed but pending a push through our buildfarm 4 - Waiting on CI Review etc done, waiting for CI to finish labels Dec 14, 2021
@sklam
sklam merged commit 23918f0 into numba:master Dec 15, 2021
@sklam sklam changed the title Testhound/fp16 support Add FP16 support for CUDA Dec 15, 2021
@gmarkall gmarkall mentioned this pull request Jan 11, 2022
1 of 7 tasks
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

5 - Ready to merge Review and testing done, is ready to merge BuildFarm Passed For PRs that have been through the buildfarm and passed CUDA CUDA related issue/PR Effort - medium Medium size effort needed

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants