Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -14,6 +14,8 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),

### Fixed

- Training on COCO-format datasets of PNG, BMP or other non-JPEG images (and of JPEG files when `simplejpeg` is not installed) loads data as fast as 1.10 again. Since 1.11.0, `CocoDetection` copied every image Pillow decoded into a NumPy array and straight back into a PIL image. On Windows 11 (Ryzen 7 7800X3D), the train DataLoader with the default 2 workers and a batch size of 8 (`RFDETRLarge`, `resolution=1120`) took up to 27% longer per batch than 1.10.1 on 4096-px BMP files and 11% longer on 4096-px PNG files; JPEG files, which `simplejpeg` decodes, were not affected. `CocoDetection` and the WebDataset shard reader now use the image Pillow decoded as is. ([#1544](https://github.com/roboflow/rf-detr/issues/1544))

- `format="coreai"` works with `coreai-torch` 0.4.3, which runs the optimization passes inside `TorchConverter.to_coreai()` and removed `AIProgram.optimize()`; with 1.11.0 a fresh `pip install "rfdetr[coreai]"` resolved 0.4.3 and every Core AI export failed with `'AIProgram' object has no attribute 'optimize'`. The `[coreai]` extra now pins `coreai-torch==0.4.3` and installs on Python 3.11 to 3.14, since `coreai-core` 1.0.0b3 ships cp314 wheels. It is declared as a uv conflict with `[tflite]`, whose `onnx2tf` pins cannot meet `coreai-core`'s `numpy>=2.3`.

- `RFDETR.from_checkpoint("checkpoint_best_total.pth")` now rebuilds the model the way it was trained. `strip_checkpoint` dropped `model_config`, so the reload fell back to class defaults with no warning: a Nano trained at `resolution=224` came back at 384 with boxes up to 193 px off, and a keypoint model trained at 552 came back at 576. `num_select` was lost the same way and `dec_layers` only logged a generic partial-load warning. The Roboflow SDK upload goes through the same call. `from_checkpoint` no longer takes `device` from the checkpoint either, so GPU-trained checkpoints load and predict on a CPU-only machine (`checkpoint_best_ema.pth` and `checkpoint_best_regular.pth` failed there before). Best-total files written by 1.11.0 and earlier can't be repaired in place; `from_checkpoint` warns about them and points at the unstripped checkpoint next to them and at `training_config.json`. ([#1533](https://github.com/roboflow/rf-detr/issues/1533))
Expand Down
16 changes: 8 additions & 8 deletions src/rfdetr/datasets/coco.py
Original file line number Diff line number Diff line change
Expand Up @@ -43,7 +43,7 @@
Resize,
)
from rfdetr.datasets.aug_configs import AUG_CONFIG
from rfdetr.datasets.io_utils import decode_image
from rfdetr.datasets.io_utils import decode_pil_image
from rfdetr.datasets.kornia_transforms import is_gpu_postprocess, resolve_backend_for_build
from rfdetr.datasets.transforms import AlbumentationsWrapper, Normalize
from rfdetr.utilities.logger import get_logger
Expand Down Expand Up @@ -306,10 +306,11 @@ def draft_size_for_transforms(
) -> int | None:
"""Return the source extent below which the transform pipeline starts losing detail.

:meth:`CocoDetection._decode_image` passes this to ``PIL.Image.draft`` so JPEG sources far larger than the training
resolution are decoded at a reduced DCT scale instead of at full size. ``draft`` never returns an image smaller
than the requested box, so the box preserves the largest direct-resize target. Scale jitter also preserves its
600-pixel pre-crop resize floor, avoiding an extra upsample after JPEG decoding.
:meth:`CocoDetection._decode_image` passes this to the decoder, which applies the reduction ``PIL.Image.draft``
would pick, so JPEG sources far larger than the training resolution are decoded at a reduced DCT scale instead of at
full size. ``draft`` never returns an image smaller than the requested box, so the box preserves the largest
direct-resize target. Scale jitter also preserves its 600-pixel pre-crop resize floor, avoiding an extra upsample
after JPEG decoding.

Two cases return ``None`` (decode at full resolution):

Expand Down Expand Up @@ -698,7 +699,7 @@ def __init__(
)

def _decode_image(self, image_id: int) -> tuple[Image.Image, tuple[float, float]]:
"""Decode one image through :func:`decode_image`, drafting when ``draft_size`` is set.
"""Decode one image through :func:`decode_pil_image`, drafting when ``draft_size`` is set.

Used instead of ``torchvision.datasets.CocoDetection._load_image``, which this class no longer calls, and
deliberately not named the same: it returns a decode scale alongside the image.
Expand All @@ -710,8 +711,7 @@ def _decode_image(self, image_id: int) -> tuple[Image.Image, tuple[float, float]
Decoded RGB image and its horizontal/vertical decode scales, both ``1.0`` when the decoder did not reduce.
"""
path = self.coco.loadImgs(image_id)[0]["file_name"]
pixels, scales = decode_image(Path(self.root) / path, self._draft_size)
return Image.fromarray(pixels), scales
return decode_pil_image(Path(self.root) / path, self._draft_size)

def __getitem__(self, idx: int) -> tuple[Any, Any]:
image_id = self.ids[idx]
Expand Down
207 changes: 159 additions & 48 deletions src/rfdetr/datasets/io_utils.py
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,20 @@
# Copyright (c) 2025 Roboflow. All Rights Reserved.
# Licensed under the Apache License, Version 2.0 [see LICENSE for details]
# ------------------------------------------------------------------------
"""Image decoding shared by the dataset readers, independent of any annotation format."""
"""Image decoding shared by the dataset readers, independent of any annotation format.

The module exposes two output-type families sharing one decoder policy: :func:`decode_image` and
:func:`decode_image_bytes` return NumPy arrays, while :func:`decode_pil_image` and :func:`decode_pil_image_bytes` return
PIL images; see below for why the split exists. Most of the array-out speedup :func:`decode_image` describes comes from
skipping Pillow's ``convert("RGB")`` and ``np.array`` copies rather than from a faster codec: the raw libjpeg-turbo
decode itself is only about 12% faster, since Pillow's wheels link the same library. Readers whose transforms take a
PIL image, ``CocoDetection`` and the WebDataset reader, use :func:`decode_pil_image` and :func:`decode_pil_image_bytes`
instead: same policy, but a Pillow-decoded image is returned as is, because copying it into an array only for the caller
to copy it back with ``Image.fromarray`` slowed their data loading on large PNG and BMP files (#1544). ``YoloDetection``
still wraps the arrays of ``_LazyYoloDetectionDataset`` with ``Image.fromarray``, as it did before ``simplejpeg``
support; that keeps its non-JPEG files (and any JPEG ``simplejpeg`` rejects) paying the same copy into an array and back
that :func:`decode_pil_image` avoids for its own callers, a scope call deferred rather than folded into this fix.
"""

from __future__ import annotations

Expand Down Expand Up @@ -82,38 +95,105 @@ def _check_decompression_bomb(width: int, height: int) -> None:
)


def _decode_with_pillow(
source: Path | IO[bytes], draft_size: int | None
) -> tuple[NDArray[np.uint8], tuple[float, float]]:
"""Decode a file or an in-memory stream through Pillow into a writeable RGB array with its decode scales.
def _decode_with_pillow(source: Path | IO[bytes], draft_size: int | None) -> tuple[Image.Image, tuple[float, float]]:
"""Decode a file or an in-memory stream through Pillow into an RGB image with its decode scales.

The fallback both entry points share: ``PIL.Image.draft`` applies the JPEG reduction when ``draft_size`` is set
(a no-op for other formats) and ``np.array`` copies the converted pixels into a buffer that outlives the closed
image.
The Pillow half of the decoder policy every entry point shares: ``PIL.Image.draft`` applies the JPEG reduction when
``draft_size`` is set (a no-op for other formats) and ``convert("RGB")`` loads the pixels into an image that
outlives the closed source.

Args:
source: Image file path, or a binary stream positioned at the start of the encoded bytes.
draft_size: Smallest extent the caller can consume without upscaling, or ``None`` for full resolution.

Returns:
Decoded ``(H, W, 3)`` uint8 RGB pixels and their horizontal/vertical decode scales, both ``1.0`` when the
decoder did not reduce.
Decoded RGB image and its horizontal/vertical decode scales, both ``1.0`` when the decoder did not reduce.

Examples:
>>> encoded = io.BytesIO()
>>> Image.fromarray(np.zeros((32, 64, 3), dtype=np.uint8)).save(encoded, format="JPEG")
>>> pixels, scales = _decode_with_pillow(io.BytesIO(encoded.getvalue()), draft_size=16)
>>> pixels.shape, scales
((16, 32, 3), (0.5, 0.5))
>>> image, scales = _decode_with_pillow(io.BytesIO(encoded.getvalue()), draft_size=16)
>>> image.mode, image.size, scales
('RGB', (32, 16), (0.5, 0.5))
"""
with Image.open(source) as image:
full_width, full_height = image.size
if draft_size is not None:
image.draft("RGB", (draft_size, draft_size))
pixels = np.array(image.convert("RGB"))
rgb_image = image.convert("RGB")
return rgb_image, (rgb_image.width / full_width, rgb_image.height / full_height)


def _decode_with_simplejpeg(
data: bytes, draft_size: int | None
) -> tuple[NDArray[np.uint8], tuple[float, float]] | None:
"""Decode JPEG bytes with ``simplejpeg``, or return ``None`` to leave them to Pillow.

``None`` when ``simplejpeg`` is not installed, when ``data`` does not start with a JPEG marker, or when
``simplejpeg`` rejects it as corrupt or unsupported, so that Pillow decodes it or raises its usual error.

Args:
data: Encoded image bytes.
draft_size: Smallest extent the caller can consume without upscaling, or ``None`` for full resolution.

Returns:
Decoded ``(H, W, 3)`` uint8 RGB pixels and their horizontal/vertical decode scales, or ``None``.

Raises:
PIL.Image.DecompressionBombError: If the image has more than twice ``PIL.Image.MAX_IMAGE_PIXELS`` pixels.

Examples:
>>> _decode_with_simplejpeg(b"not a jpeg", draft_size=None) is None
True
"""
if simplejpeg is None or not data.startswith(_JPEG_SOI):
return None
try:
header = simplejpeg.decode_jpeg_header(data)
full_height, full_width = int(header[0]), int(header[1])
# Pillow's guard on public names, so the limit and its warn/raise tiers stay those of ``Image.open``.
_check_decompression_bomb(full_width, full_height)
reduction = _jpeg_draft_reduction(full_width, full_height, draft_size)
pixels = simplejpeg.decode_jpeg(
data,
colorspace="RGB",
min_height=-(-full_height // reduction),
min_width=-(-full_width // reduction),
)
except ValueError:
return None # corrupt or unsupported JPEG: let Pillow decode it or raise its usual error
return pixels, (pixels.shape[1] / full_width, pixels.shape[0] / full_height)


def _read_simplejpeg_candidate(path: Path) -> bytes | None:
"""Return the file's bytes when ``simplejpeg`` is installed and the file starts with a JPEG marker, else ``None``.

The marker decides, not the extension, so a mislabelled file still reaches the decoder that can read it. Without
``simplejpeg`` nothing is read here, and Pillow streams the file itself.

Args:
path: Image file to inspect.

Returns:
The encoded bytes for :func:`_decode_with_simplejpeg`, or ``None`` when the file is Pillow's to decode.

Examples:
>>> import tempfile
>>> with tempfile.TemporaryDirectory() as tmp:
... png_path = Path(tmp) / "img.png"
... Image.new("RGB", (4, 4)).save(png_path)
... _read_simplejpeg_candidate(png_path) is None
True
"""
if simplejpeg is None:
return None
with path.open("rb") as encoded:
if encoded.read(len(_JPEG_SOI)) != _JPEG_SOI:
return None
encoded.seek(0)
return encoded.read()


def decode_image(path: Path, draft_size: int | None = None) -> tuple[NDArray[np.uint8], tuple[float, float]]:
"""Decode an image file to RGB, optionally downscaling during the JPEG discrete cosine transform (DCT) decode.

Expand All @@ -127,14 +207,7 @@ def decode_image(path: Path, draft_size: int | None = None) -> tuple[NDArray[np.

Returning an array rather than a PIL image lets a caller that wants an array skip a round trip, which is where the
speedup lands: 1.3-1.8x over Pillow at the decode stage for array consumers such as ``_LazyYoloDetectionDataset``,
varying with image size and with how much high-frequency detail the JPEG carries. Most of that array-out gain comes
from skipping Pillow's ``convert("RGB")`` and ``np.array`` copies rather than from a faster codec: the raw
libjpeg-turbo decode itself is only about 12% faster, since Pillow's wheels link the same library. Callers needing
a PIL image wrap the result with ``Image.fromarray`` (``YoloDetection``, ``CocoDetection``, the WebDataset reader);
that wrap costs about what the faster decode saves, so the net effect there is hardware-dependent: ``YoloDetection``
keeps a smaller end-to-end gain because its previous path copied through ``np.array`` too, while other PIL-out
consumers may see a slight regression on some platforms until the CPU pipeline consumes arrays directly, and take
the shared decoder policy rather than a speedup.
varying with image size and with how much high-frequency detail the JPEG carries.

When ``draft_size`` is set, both decoders apply the same power-of-two reduction ``PIL.Image.draft`` would choose to
keep the image at least ``draft_size`` on both axes; it is a no-op for non-JPEG files.
Expand All @@ -150,20 +223,18 @@ def decode_image(path: Path, draft_size: int | None = None) -> tuple[NDArray[np.
Raises:
PIL.Image.DecompressionBombError: If the image has more than twice ``PIL.Image.MAX_IMAGE_PIXELS`` pixels.
"""
if simplejpeg is not None:
with path.open("rb") as encoded:
if encoded.read(len(_JPEG_SOI)) == _JPEG_SOI:
encoded.seek(0)
return decode_image_bytes(encoded.read(), draft_size)

return _decode_with_pillow(path, draft_size)
data = _read_simplejpeg_candidate(path)
if data is not None:
return decode_image_bytes(data, draft_size)
image, scales = _decode_with_pillow(path, draft_size)
return np.array(image), scales


def decode_image_bytes(data: bytes, draft_size: int | None = None) -> tuple[NDArray[np.uint8], tuple[float, float]]:
"""Decode encoded image bytes to RGB, optionally downscaling during the JPEG DCT (discrete cosine transform) decode.

Holds the decoder policy :func:`decode_image` applies; readers that already have the encoded bytes in memory, such
as the WebDataset loader reading a shard member, call this directly instead of writing the file out first.
Holds the decoder policy :func:`decode_image` applies to encoded bytes already in memory; :func:`decode_image`
routes JPEG files through it, and :func:`decode_pil_image_bytes` is its counterpart that returns a PIL image.

Args:
data: Encoded image bytes.
Expand All @@ -176,21 +247,61 @@ def decode_image_bytes(data: bytes, draft_size: int | None = None) -> tuple[NDAr
Raises:
PIL.Image.DecompressionBombError: If the image has more than twice ``PIL.Image.MAX_IMAGE_PIXELS`` pixels.
"""
if simplejpeg is not None and data.startswith(_JPEG_SOI):
try:
header = simplejpeg.decode_jpeg_header(data)
full_height, full_width = int(header[0]), int(header[1])
# Pillow's guard on public names, so the limit and its warn/raise tiers stay those of ``Image.open``.
_check_decompression_bomb(full_width, full_height)
reduction = _jpeg_draft_reduction(full_width, full_height, draft_size)
pixels = simplejpeg.decode_jpeg(
data,
colorspace="RGB",
min_height=-(-full_height // reduction),
min_width=-(-full_width // reduction),
)
except ValueError:
pass # corrupt or unsupported JPEG: let Pillow decode it or raise its usual error
else:
return pixels, (pixels.shape[1] / full_width, pixels.shape[0] / full_height)
decoded = _decode_with_simplejpeg(data, draft_size)
if decoded is not None:
return decoded
image, scales = _decode_with_pillow(io.BytesIO(data), draft_size)
return np.array(image), scales


def decode_pil_image(path: Path, draft_size: int | None = None) -> tuple[Image.Image, tuple[float, float]]:
"""Decode an image file to an RGB PIL image under the same decoder policy as :func:`decode_image`.

For readers whose transforms take a PIL image, such as ``CocoDetection``. A JPEG that ``simplejpeg`` decodes is
wrapped with ``Image.fromarray``; everything Pillow decodes is returned as Pillow produced it, without the copy
into an array and back that :func:`decode_image` plus ``Image.fromarray`` would make. Pixels and decode scales are
the ones :func:`decode_image` returns for the same file. The default Albumentations CPU training wrapper still
round-trips a decoded JPEG through its own ``np.array``/``Image.fromarray`` pair on top of this, independent of
this policy. The returned image's ``info`` dict is populated only when Pillow decoded it; a JPEG that
``simplejpeg`` decoded goes through ``Image.fromarray``, which always starts with an empty ``info``, so ``info``
is not part of this function's contract.

Args:
path: Image file to decode.
draft_size: Smallest extent the caller can consume without upscaling, or ``None`` for full resolution.

Returns:
Decoded RGB image and its horizontal/vertical decode scales, both ``1.0`` when the decoder did not reduce.

Raises:
PIL.Image.DecompressionBombError: If the image has more than twice ``PIL.Image.MAX_IMAGE_PIXELS`` pixels.
"""
data = _read_simplejpeg_candidate(path)
if data is not None:
return decode_pil_image_bytes(data, draft_size)
return _decode_with_pillow(path, draft_size)


def decode_pil_image_bytes(data: bytes, draft_size: int | None = None) -> tuple[Image.Image, tuple[float, float]]:
"""Decode encoded image bytes to an RGB PIL image under the same decoder policy as :func:`decode_image_bytes`.

The in-memory counterpart of :func:`decode_pil_image`, for readers such as the WebDataset loader that hold a shard
member's bytes and feed PIL-based transforms. As in :func:`decode_pil_image`, the returned image's ``info`` dict
is populated only when Pillow decoded it, not when ``simplejpeg`` did (``Image.fromarray`` starts with an empty
``info``); ``info`` is not part of this function's contract.

Args:
data: Encoded image bytes.
draft_size: Smallest extent the caller can consume without upscaling, or ``None`` for full resolution.

Returns:
Decoded RGB image and its horizontal/vertical decode scales, both ``1.0`` when the decoder did not reduce.

Raises:
PIL.Image.DecompressionBombError: If the image has more than twice ``PIL.Image.MAX_IMAGE_PIXELS`` pixels.
"""
decoded = _decode_with_simplejpeg(data, draft_size)
if decoded is not None:
pixels, scales = decoded
return Image.fromarray(pixels), scales
return _decode_with_pillow(io.BytesIO(data), draft_size)
Loading
Loading