Skip to content

fix: resample multiscale labels with nearest neighbour in transform() - #1267

Open
Sizerta wants to merge 1 commit into
scverse:mainfrom
Sizerta:fix-multiscale-labels-nearest
Open

Sizerta wants to merge 1 commit into
scverse:mainfrom
Sizerta:fix-multiscale-labels-nearest

Conversation

@Sizerta

@Sizerta Sizerta commented Oct 2, 2026 •

Copy link
Copy Markdown

Closes #1202

Hello there. Thanks for the clear report and the reproduction script, they made this easy to track down. I can confirm the bug on current main: starting from a mask that contains only the ids 0, 1 and 50, rotating the multiscale labels by 30° gave 44 distinct ids at scale0 and 33 at scale1, and Scale(1.5) produced id 17. Single-scale labels were correct.

Cause

Segmentation labels are categorical data: each value names an instance rather than measuring an intensity, so there is no meaningful value between cell 1 and cell 50. Linear interpolation suits intensities, but on labels it produces weighted averages of neighbouring ids along every boundary. These either match no row of the annotating table, or coincide with the id of an unrelated cell, which then silently propagates into anything computed per instance, such as aggregation or table joins. Nearest-neighbour resampling (order=0) is the appropriate choice, since every output pixel then copies an id that exists in the input.

In the DataTree overload of transform(), labels were resampled with only {"prefilter": False}, so dask_image.ndinterp.affine_transform used its default order=1. The DataArray overload already uses order=0.

Fix

I added order=0 for labels in the DataTree overload and replaced the # TODO: this should work, test better comment there with a short explanation. transform(sdata, ...) and transform_to_coordinate_system() use the same overload, so they're covered too.

Tests

The existing test_transform_raster compares extents only, so the label values weren't checked. The new tests assert that, at every resolution level, the set of ids is a subset of the input ids (nothing is invented) and still contains both cells (so the check can't pass trivially). They cover:

  • rotation and scaling
  • single-scale and multiscale labels
  • the element on its own and inside a SpatialData object
  • 3D labels (Labels3DModel)

Without the fix, the 5 multiscale cases fail; with it, all 10 pass. The full suite goes from 1373 to 1383 passing tests, with no other changes.

Not included

The issue also notes that single-scale images are resampled with nearest neighbour, while multiscale images are linearly interpolated. Since choosing one policy for images (and possibly exposing order) is a design decision, I left it for #725

@codecov

codecov Bot commented Oct 2, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 91.40%. Comparing base (ea93a37) to head (adea5d7).

Additional details and impacted files
@@           Coverage Diff           @@
##             main    #1267   +/-   ##
=======================================
  Coverage   91.40%   91.40%           
=======================================
  Files          53       53           
  Lines        8381     8381           
=======================================
  Hits         7661     7661           
  Misses        720      720           
Files with missing lines Coverage Δ
src/spatialdata/_core/operations/transform.py 92.00% <100.00%> (ø)
🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

transform() of multiscale labels uses linear interpolation and invents label ids

1 participant