Skip to content

Fancy-index assignment / add.at into zero-length array corrupts CUDA context (Warp Illegal Address) instead of raising IndexError #10245

Description

@btvn-hy-v

Summary

Fancy-index assignment / add.at into a zero-length array corrupts the CUDA context with a sticky cudaErrorIllegalAddress instead of raising IndexError. NumPy raises for the same code.

Per the documented difference from NumPy, CuPy wraps out-of-bounds integer indices instead of raising (https://docs.cupy.dev/en/latest/user_guide/difference.html#out-of-bounds-indices). That wrap is implemented as index % n against the target length — which is a division by zero when n == 0. Every index in a non-empty index list is necessarily out of bounds for a zero-length target, so there is no valid destination for any element: the only sane behaviors are raising IndexError or skipping all writes. Instead the kernel computes a garbage address and faults, and the sticky error kills every subsequent CUDA call in the process.

Reproduction

import cupy as cp

a = cp.zeros(0, dtype=cp.bool_)
a[cp.array([6, 5, 1], dtype=cp.int32)] = True   # cupy_scatter_update
cp.cuda.runtime.deviceSynchronize()
# cupy_backends.cuda.api.runtime.CUDARuntimeError: cudaErrorIllegalAddress:
# an illegal memory access was encountered

The same happens with add.at:

b = cp.zeros(0, dtype=cp.float64)
cp.add.at(b, cp.array([6, 5, 1], dtype=cp.int32), 1)   # cupy_scatter_add
cp.cuda.runtime.deviceSynchronize()
# CUDARuntimeError: cudaErrorIllegalAddress

NumPy for comparison:

import numpy as np
x = np.zeros(0)
x[[6, 5, 1]] = True
# IndexError: index 6 is out of bounds for axis 0 with size 0

Environment

  • CuPy 14.1.1 (wheel, CUDA 12.x)
  • NVIDIA B200 (sm_100), driver 580.x
  • Reproduced with a driver GPU core dump: faulting kernel is cupy_scatter_update<<<(1,1,1),(3,1,1)>>>, exception Warp Illegal Address — block dim equals the index count, confirming the scatter path

Impact

A single stray setitem/add.at against an empty array permanently poisons the CUDA context of the whole process (every later CUDA API returns cudaErrorIllegalAddress). This took down a production query server: the fault surfaced asynchronously, error handling could not recover the process, and the sticky context required a restart.

Suggested fix

Guard the empty-target case in the scatter kernels' host-side launch path (or the indexing entry points ndarray.__setitem__ / ndarray.add.at): if the target axis has length 0 and the index list is non-empty, raise IndexError (NumPy semantics). Alternatively skip out-of-bounds/empty-target writes entirely — but the current modulo-by-zero should not reach the device.

Happy to provide the GPU core dump or the full server-side trace if useful.

Root cause in source

_create_scatter_kernel in cupy/_core/_routines_indexing.pyx (backs all scatter ops: __setitem__, add.at, etc.) wraps indices as indices % adim with no adim == 0 guard — on GPU the integer division by zero yields an undefined value, so the store hits a garbage address. Sibling paths already handle this case (_prepare_multiple_array_indexing guards if a_shape_i != 0; _take raises IndexError for empty axes), so a host-side check in _scatter_op_single matching NumPy's IndexError should cover it.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions