Skip to content

BUG: Highly degraded precision for tanh for arrays of float16 on Sapphire Rapids - Regression in 2.3.0 #30821

Description

@Flamefire

Describe the issue:

I ran into an issue I traced to numpy, that was introduced in 2.3.0, 2.2.6 works fine.
I only observed it on Intel Sapphire Rapids, but not AMD Zen3 CPUs

Calculating tanh of Float16 values has a large error when processing arrays. Most notably it saturates at 0.9985 instead of 1

Reproduce the code example:

$ python -c 'import numpy as np; print(np.tanh(np.float16(1000)), np.tanh(np.array([1000], dtype=np.float16)), np.tanh(np.array([1000], dtype=np.float32)))'
1.0 [0.9985] [1.]

Python and NumPy Versions:

2.4.2, but observed since 2.3.0, NOT with 2.2.6

Python 3.11.5 (main, Dec 17 2024, 11:35:51) [GCC 13.2.0], but doesn't matter

Runtime Environment:

[{'numpy_version': '2.4.2',
'python': '3.11.5 (main, Dec 17 2024, 11:35:51) [GCC 13.2.0]',
'uname': uname_result(system='Linux', node='login1', release='5.14.0-570.49.1.el9_6.x86_64', version='#1 SMP PREEMPT_DYNAMIC Fri Oct 3 15:42:32 UTC 2025', machine='x86_64')},
{'simd_extensions': {'baseline': ['X86_V2'],
'found': ['X86_V3', 'X86_V4', 'AVX512_ICL', 'AVX512_SPR'],
'not_found': []}},
{'ignore_floating_point_errors_in_matmul': False}]

How does this issue affect you or how did you find it:

E.g. the mish operation for ML/AI relies on tanh, specifically: x * tanh(softmax(x)). With this issue results are wildly off (due to multiplication with x amplifying the error)

Found in the PyTorch test suite where numpy is used as the reference

Possibly related: #25934 1fa958d

Activity

  1. seberg commented on Feb 11, 2026

    @seberg
    Member
  2. BiradarSiddhant02 commented on Feb 18, 2026

    @BiradarSiddhant02

    An issue in pytorch seems to be similar.
    pytorch/pytorch#173757

    The above linked issue is more or a performance issue but worth looking at.

  3. r-devulap commented on Feb 18, 2026

    @r-devulap
    Member

    This is coming from the 4 ULP SVML implementation of FP16 tanh coming from https://github.com/numpy/SVML/blob/main/linux/avx512/svml_z0_atanh_h_la.s which only works on SPR.
    0.9985 and 1 are 3 ULP away, so this is within the 4 ULP error range.

  4. BiradarSiddhant02 commented on Feb 18, 2026

    @BiradarSiddhant02

    This is coming from the 4 ULP SVML implementation of FP16 tanh coming from https://github.com/numpy/SVML/blob/main/linux/avx512/svml_z0_atanh_h_la.s which only works on SPR. 0.9985 and 1 are 3 ULPS away, so this is within the 4 ULP error range.

    Ah ok makes sense!

  5. Flamefire commented on Feb 19, 2026

    @Flamefire
    ContributorAuthor

    Can that be solved? Especially the saturation at less than 1 might cause issues with some applications, e.g. ML stuff could expect tanh(x) == 1 for large x, although, of course, it is mathematically only a limit

  6. seberg commented on Feb 19, 2026

    @seberg
    Member

    We probably have to have a discussion whether 4ULP is sensible here. There was a bit of an argument always that for these low precision types the users tend to want a fast tradeoff, but to some degrees I think these opinions on speed vs. precision don't have a super good foundation.
    (Of course it is very hard to actually put them on a good foundation.)

    Beyond ULP as you note there is also always the problem that certain behaviors can be very surprising. We had the issue with 64bit sin/cos which were probably precise enough but code failed because it wasn't monotonous enough (1 or 2 ULP jitters where the numerical value was ~1)...

    If we decide against 4ULP for lower precision fps too, then there may still be a path to enable "fast but low precision". I am a bit wary about these low precision versions in NumPy, but I think there are also strong opinions that NumPy should not sacrifice speed too much.

    Do we have a higher precision fp16 version and know the speed trade-off?

  7. BiradarSiddhant02 commented on Feb 19, 2026

    @BiradarSiddhant02

    I'm planning to put together a detailed benchmark and report comparing different fp16/bf16 math implementations. I may open a repo and link it here.
    In the meantime, this is worth reading: https://www.pythontutorials.net/blog/float16-is-much-slower-than-float32-and-float64-in-numpy/
    As that article shows, the speed argument for fp16 is already compromised at the instruction level on CPUs, the emulation overhead from converting to/from fp32 erases any throughput advantage. Given that, there's little remaining justification for sacrificing precision in fp16 math functions. If we're not actually getting the speed, we may as well get the accuracy.

  8. seberg commented on Feb 19, 2026

    @seberg
    Member

    I doubt that is relevant. These changes here are precisely because float16 isn't just done in float32 anymore in practice which was the case before this change.

  9. r-devulap commented on Feb 20, 2026

    @r-devulap
    Member

    Do we have a higher precision fp16 version and know the speed trade-off?

    It doesn’t look like it. The repository only includes the 4‑ULP variants of the float16 functions: https://github.com/numpy/SVML/tree/main/linux/avx512. I’m no longer at Intel, but I can try reaching out to a few folks to see whether the 1‑ULP versions can be open‑sourced. If it happens, it may take some time before they make it upstream.
    In the meantime, should we consider reverting the current tanh implementation?

  10. added this to the 2.4.3 release milestone on Feb 20, 2026
  11. seberg commented on Feb 20, 2026

    @seberg
    Member

    In the meantime, should we consider reverting the current tanh implementation?

    Yeah probably. I have to see what other functions we have with low precision here.

  12. r-devulap commented on Feb 20, 2026

    @r-devulap
    Member

    See #23351. The functions and corresponding errors are listed at the top.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions