Summary
Fancy-index assignment / add.at into a zero-length array corrupts the CUDA context with a sticky cudaErrorIllegalAddress instead of raising IndexError. NumPy raises for the same code.
Per the documented difference from NumPy, CuPy wraps out-of-bounds integer indices instead of raising (https://docs.cupy.dev/en/latest/user_guide/difference.html#out-of-bounds-indices). That wrap is implemented as index % n against the target length — which is a division by zero when n == 0. Every index in a non-empty index list is necessarily out of bounds for a zero-length target, so there is no valid destination for any element: the only sane behaviors are raising IndexError or skipping all writes. Instead the kernel computes a garbage address and faults, and the sticky error kills every subsequent CUDA call in the process.
Reproduction
import cupy as cp
a = cp.zeros(0, dtype=cp.bool_)
a[cp.array([6, 5, 1], dtype=cp.int32)] = True # cupy_scatter_update
cp.cuda.runtime.deviceSynchronize()
# cupy_backends.cuda.api.runtime.CUDARuntimeError: cudaErrorIllegalAddress:
# an illegal memory access was encountered
The same happens with add.at:
b = cp.zeros(0, dtype=cp.float64)
cp.add.at(b, cp.array([6, 5, 1], dtype=cp.int32), 1) # cupy_scatter_add
cp.cuda.runtime.deviceSynchronize()
# CUDARuntimeError: cudaErrorIllegalAddress
NumPy for comparison:
import numpy as np
x = np.zeros(0)
x[[6, 5, 1]] = True
# IndexError: index 6 is out of bounds for axis 0 with size 0
Environment
- CuPy 14.1.1 (wheel, CUDA 12.x)
- NVIDIA B200 (sm_100), driver 580.x
- Reproduced with a driver GPU core dump: faulting kernel is
cupy_scatter_update<<<(1,1,1),(3,1,1)>>>, exception Warp Illegal Address — block dim equals the index count, confirming the scatter path
Impact
A single stray setitem/add.at against an empty array permanently poisons the CUDA context of the whole process (every later CUDA API returns cudaErrorIllegalAddress). This took down a production query server: the fault surfaced asynchronously, error handling could not recover the process, and the sticky context required a restart.
Suggested fix
Guard the empty-target case in the scatter kernels' host-side launch path (or the indexing entry points ndarray.__setitem__ / ndarray.add.at): if the target axis has length 0 and the index list is non-empty, raise IndexError (NumPy semantics). Alternatively skip out-of-bounds/empty-target writes entirely — but the current modulo-by-zero should not reach the device.
Happy to provide the GPU core dump or the full server-side trace if useful.
Root cause in source
_create_scatter_kernel in cupy/_core/_routines_indexing.pyx (backs all scatter ops: __setitem__, add.at, etc.) wraps indices as indices % adim with no adim == 0 guard — on GPU the integer division by zero yields an undefined value, so the store hits a garbage address. Sibling paths already handle this case (_prepare_multiple_array_indexing guards if a_shape_i != 0; _take raises IndexError for empty axes), so a host-side check in _scatter_op_single matching NumPy's IndexError should cover it.
Summary
Fancy-index assignment /
add.atinto a zero-length array corrupts the CUDA context with a stickycudaErrorIllegalAddressinstead of raisingIndexError. NumPy raises for the same code.Per the documented difference from NumPy, CuPy wraps out-of-bounds integer indices instead of raising (
https://docs.cupy.dev/en/latest/user_guide/difference.html#out-of-bounds-indices). That wrap is implemented asindex % nagainst the target length — which is a division by zero whenn == 0. Every index in a non-empty index list is necessarily out of bounds for a zero-length target, so there is no valid destination for any element: the only sane behaviors are raisingIndexErroror skipping all writes. Instead the kernel computes a garbage address and faults, and the sticky error kills every subsequent CUDA call in the process.Reproduction
The same happens with
add.at:NumPy for comparison:
Environment
cupy_scatter_update<<<(1,1,1),(3,1,1)>>>, exceptionWarp Illegal Address— block dim equals the index count, confirming the scatter pathImpact
A single stray setitem/
add.atagainst an empty array permanently poisons the CUDA context of the whole process (every later CUDA API returnscudaErrorIllegalAddress). This took down a production query server: the fault surfaced asynchronously, error handling could not recover the process, and the sticky context required a restart.Suggested fix
Guard the empty-target case in the scatter kernels' host-side launch path (or the indexing entry points
ndarray.__setitem__/ndarray.add.at): if the target axis has length 0 and the index list is non-empty, raiseIndexError(NumPy semantics). Alternatively skip out-of-bounds/empty-target writes entirely — but the current modulo-by-zero should not reach the device.Happy to provide the GPU core dump or the full server-side trace if useful.
Root cause in source
_create_scatter_kernelincupy/_core/_routines_indexing.pyx(backs all scatter ops:__setitem__,add.at, etc.) wraps indices asindices % adimwith noadim == 0guard — on GPU the integer division by zero yields an undefined value, so the store hits a garbage address. Sibling paths already handle this case (_prepare_multiple_array_indexingguardsif a_shape_i != 0;_takeraisesIndexErrorfor empty axes), so a host-side check in_scatter_op_singlematching NumPy'sIndexErrorshould cover it.