Skip to content

relax CUDASIM test fp16 intrinsic comparison tolerance. - #10822

Merged
swap357 merged 2 commits into
numba:mainfrom
swap357:fix/10807-cudasim-fp16-tolerance
Sep 8, 2026
Merged

swap357 merged 2 commits into
numba:mainfrom
swap357:fix/10807-cudasim-fp16-tolerance

Conversation

@swap357

@swap357 swap357 commented Sep 8, 2026

Copy link
Copy Markdown
Contributor

Closes #10807

@swap357
swap357 requested a review from gmarkall as a code owner September 8, 2026 17:36
@swap357 swap357 added 3 - Ready for Review skip_release_notes Skip towncrier requirement labels Sep 8, 2026

@esc esc left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

1e-3 is fine.

The test at:

https://github.com/numba/numba/actions/runs/33952384424/job/101269798444)

shows the following for the log10 usecase:

Mismatched elements: 1 / 32 (3.12%)
Max absolute difference among violations: 0.003906
Max relative difference among violations: 0.0008197
 ACTUAL: array([4.797, 4.52 , 4.086, 3.715, 4.51 , 4.7  , 4.64 , 3.893, 4.715,
       4.332, 4.516, 4.312, 4.69 , 3.889, 4.64 , 4.496, 4.56 , 4.504,
       4.77 , 4.74 , 4.336, 4.656, 4.555, 4.715, 4.81 , 3.926, 4.3  ,
       4.574, 4.758, 4.36 , 4.695, 4.434], dtype=float16)
 DESIRED: array([4.797, 4.52 , 4.086, 3.715, 4.51 , 4.7  , 4.64 , 3.893, 4.715,
       4.332, 4.516, 4.312, 4.69 , 3.889, 4.64 , 4.496, 4.56 , 4.504,
       4.766, 4.74 , 4.336, 4.656, 4.555, 4.715, 4.81 , 3.926, 4.3  ,
       4.574, 4.758, 4.36 , 4.695, 4.434], dtype=float16)

If you put that in a Python program:

import numpy as np

# ACTUAL vs DESIRED from the failing CI job (fn=log10), numba#10807
actual  = np.array([4.797, 4.52, 4.086, 3.715, 4.51, 4.7, 4.64, 3.893, 4.715,
                    4.332, 4.516, 4.312, 4.69, 3.889, 4.64, 4.496, 4.56, 4.504,
                    4.77, 4.74, 4.336, 4.656, 4.555, 4.715, 4.81, 3.926, 4.3,
                    4.574, 4.758, 4.36, 4.695, 4.434], dtype=np.float16)
desired = np.array([4.797, 4.52, 4.086, 3.715, 4.51, 4.7, 4.64, 3.893, 4.715,
                    4.332, 4.516, 4.312, 4.69, 3.889, 4.64, 4.496, 4.56, 4.504,
                    4.766, 4.74, 4.336, 4.656, 4.555, 4.715, 4.81, 3.926, 4.3,
                    4.574, 4.758, 4.36, 4.695, 4.434], dtype=np.float16)

for rtol in (1e-7, 1e-3):
    try:
        np.testing.assert_allclose(actual, desired, rtol=rtol)
        print(f"rtol={rtol}: PASS")
    except AssertionError:
        print(f"rtol={rtol}: FAIL")

Output (NumPy 2.3.5):

rtol=1e-07: FAIL
rtol=0.001: PASS

So yes, 1e-3 is fine.

@swap357
swap357 merged commit 7118415 into numba:main Sep 8, 2026
24 checks passed
@esc esc added 5 - Ready to merge Review and testing done, is ready to merge and removed 3 - Ready for Review labels Sep 11, 2026
@esc esc added this to the 0.68.0rc1 milestone Sep 11, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

5 - Ready to merge Review and testing done, is ready to merge skip_release_notes Skip towncrier requirement

Projects

None yet

Development

Successfully merging this pull request may close these issues.

linux-64 conda CI fails on NumPy 2.3: CUDA-sim test_fp16_intrinsics_common

2 participants