The current implementation of both ufuncs and gufuncs (those with core-dimensions and an inner-loop signature) use the NPY_ITER_OVERLAP_ASSUME_ELEMENTWISE flag to prevent allocating buffer memory if it can be determined that:
- the data pointers of all overlapping operands are equal
- the strides and dimensions are equivalent
- the dtypes are equal.
solve_may_have_internal_overlap() for single-byte overlap returns `0
Let's call this element-wise aliasing, since it is intended for elementwise ufuncs like np.sin.
For all other cases, output ndarrays will use writeback semantics to allocate temporary memory.
There should be a point in the gufunc call that a gufunc can say "element-wise aliasing is OK", or "leave all aliasing to the inner loop" or "always copy-on-any-overlap". We need a flag to indicate these (and maybe other, like contiguous) strategies. See also PR #11381 (closed) which proposed unilaterally changing the default.
The current implementation of both ufuncs and gufuncs (those with core-dimensions and an inner-loop signature) use the
NPY_ITER_OVERLAP_ASSUME_ELEMENTWISEflag to prevent allocating buffer memory if it can be determined that:solve_may_have_internal_overlap()for single-byte overlap returns `0Let's call this element-wise aliasing, since it is intended for elementwise ufuncs like
np.sin.For all other cases, output ndarrays will use writeback semantics to allocate temporary memory.
There should be a point in the gufunc call that a gufunc can say "element-wise aliasing is OK", or "leave all aliasing to the inner loop" or "always copy-on-any-overlap". We need a flag to indicate these (and maybe other, like contiguous) strategies. See also PR #11381 (closed) which proposed unilaterally changing the default.