Repository navigation
BUG: Highly degraded precision for tanh for arrays of float16 on Sapphire Rapids - Regression in 2.3.0 #30821
Description
Activity
CC @sterrettm2, @Mousius.
An issue in pytorch seems to be similar.
pytorch/pytorch#173757The above linked issue is more or a performance issue but worth looking at.
This is coming from the 4 ULP SVML implementation of FP16
tanhcoming from https://github.com/numpy/SVML/blob/main/linux/avx512/svml_z0_atanh_h_la.s which only works on SPR.
0.9985 and 1 are 3 ULP away, so this is within the 4 ULP error range.This is coming from the 4 ULP SVML implementation of FP16
tanhcoming from https://github.com/numpy/SVML/blob/main/linux/avx512/svml_z0_atanh_h_la.s which only works on SPR. 0.9985 and 1 are 3 ULPS away, so this is within the 4 ULP error range.Ah ok makes sense!
Can that be solved? Especially the saturation at less than 1 might cause issues with some applications, e.g. ML stuff could expect
tanh(x) == 1for largex, although, of course, it is mathematically only a limitWe probably have to have a discussion whether 4ULP is sensible here. There was a bit of an argument always that for these low precision types the users tend to want a fast tradeoff, but to some degrees I think these opinions on speed vs. precision don't have a super good foundation.
(Of course it is very hard to actually put them on a good foundation.)Beyond ULP as you note there is also always the problem that certain behaviors can be very surprising. We had the issue with 64bit sin/cos which were probably precise enough but code failed because it wasn't monotonous enough (1 or 2 ULP jitters where the numerical value was ~1)...
If we decide against 4ULP for lower precision fps too, then there may still be a path to enable "fast but low precision". I am a bit wary about these low precision versions in NumPy, but I think there are also strong opinions that NumPy should not sacrifice speed too much.
Do we have a higher precision fp16 version and know the speed trade-off?
I'm planning to put together a detailed benchmark and report comparing different fp16/bf16 math implementations. I may open a repo and link it here.
In the meantime, this is worth reading: https://www.pythontutorials.net/blog/float16-is-much-slower-than-float32-and-float64-in-numpy/
As that article shows, the speed argument for fp16 is already compromised at the instruction level on CPUs, the emulation overhead from converting to/from fp32 erases any throughput advantage. Given that, there's little remaining justification for sacrificing precision in fp16 math functions. If we're not actually getting the speed, we may as well get the accuracy.I doubt that is relevant. These changes here are precisely because float16 isn't just done in float32 anymore in practice which was the case before this change.
Do we have a higher precision fp16 version and know the speed trade-off?
It doesn’t look like it. The repository only includes the 4‑ULP variants of the float16 functions: https://github.com/numpy/SVML/tree/main/linux/avx512. I’m no longer at Intel, but I can try reaching out to a few folks to see whether the 1‑ULP versions can be open‑sourced. If it happens, it may take some time before they make it upstream.
In the meantime, should we consider reverting the current tanh implementation?In the meantime, should we consider reverting the current tanh implementation?
Yeah probably. I have to see what other functions we have with low precision here.
See #23351. The functions and corresponding errors are listed at the top.
Reacted by Sebastian Berg
Describe the issue:
I ran into an issue I traced to numpy, that was introduced in 2.3.0, 2.2.6 works fine.
I only observed it on Intel Sapphire Rapids, but not AMD Zen3 CPUs
Calculating
tanhof Float16 values has a large error when processing arrays. Most notably it saturates at 0.9985 instead of 1Reproduce the code example:
Python and NumPy Versions:
2.4.2, but observed since 2.3.0, NOT with 2.2.6
Python 3.11.5 (main, Dec 17 2024, 11:35:51) [GCC 13.2.0], but doesn't matter
Runtime Environment:
[{'numpy_version': '2.4.2',
'python': '3.11.5 (main, Dec 17 2024, 11:35:51) [GCC 13.2.0]',
'uname': uname_result(system='Linux', node='login1', release='5.14.0-570.49.1.el9_6.x86_64', version='#1 SMP PREEMPT_DYNAMIC Fri Oct 3 15:42:32 UTC 2025', machine='x86_64')},
{'simd_extensions': {'baseline': ['X86_V2'],
'found': ['X86_V3', 'X86_V4', 'AVX512_ICL', 'AVX512_SPR'],
'not_found': []}},
{'ignore_floating_point_errors_in_matmul': False}]
How does this issue affect you or how did you find it:
E.g. the
mishoperation for ML/AI relies ontanh, specifically:x * tanh(softmax(x)). With this issue results are wildly off (due to multiplication withxamplifying the error)Found in the PyTorch test suite where numpy is used as the reference
Possibly related: #25934 1fa958d