Skip to content

immich_server leaks memory per asset while machine learning server runs #31488

Description

@donkeyteethUX

I have searched the existing issues, both open and closed, to make sure this is not a duplicate report.

  • Yes

The bug

AI disclaimer: this is written up by claude and lightly edited by me for clarity. Hopefully that's appropriate.

Fwiw, the clanker is pretty confident it knows where the memory leak is. See "Additional Information".


While re-running Smart Search for the whole library after switching CLIP models, the immich_server container's memory grew with every asset processed and was never released. On an 8 GB Raspberry Pi 5 the server process reached 6.3 GB resident after about 18,000 assets and the kernel OOM-killed it.

Measured growth is roughly 300–600 KB per asset, which is about the size of one preview JPEG. Growth tracked the number of assets processed, not the job concurrency: concurrency 6 and 12 leaked about the same amount per asset, just faster or slower. The same thing happens with the OCR job. Memory is only freed by restarting the container.

The machine learning container was on a separate machine (http://<laptop>:3003 as the only ML URL). I have not tested whether it also happens with the default local ML container, but the code path that sends the image is the same.

To get through the library I ran a loop that restarts immich_server whenever its RSS passes 3 GB. That fired every 3–4 minutes for the whole run (14 restarts for ~118k assets). Between restarts the job resumed fine, so nothing is wrong with the results, only with memory.

The OS that Immich Server is running on

Debian GNU/Linux 12 (bookworm), Raspberry Pi OS, kernel 6.12.34+rpt-rpi-2712, Raspberry Pi 5 Model B, 8 GB RAM. Docker 28.3.3, Compose 2.39.1.

Version of Immich Server

v3.2.0

Version of Immich Mobile App

n/a

Platform with the issue

  • Server
  • Web
  • Mobile

Device make and model

No response

Your docker-compose.yml content

Unmodified release docker-compose.yml for v3.2.0 (server, machine-learning, redis/valkey, postgres). The local immich-machine-learning service was still running but not used; the only ML URL configured in Administration → Machine Learning Settings was the remote one.

Your .env content

UPLOAD_LOCATION=./library
IMMICH_VERSION=release
DB_PASSWORD=<redacted>
DB_HOSTNAME=immich_postgres
DB_USERNAME=postgres
DB_DATABASE_NAME=immich
REDIS_HOSTNAME=immich_redis

Reproduction steps

  1. Library with a large number of images (mine: ~118k eligible assets).
  2. Optionally run the machine-learning container on another machine and set its URL as the ML URL. (Probably not required, see above.)
  3. Administration → Settings → Machine Learning → change the CLIP model (I went from ViT-B-32__openai to ViT-B-32-SigLIP2-256__webli), or just run Jobs → Smart Search → All.
  4. Watch docker stats immich_server or the RSS of the immich process. It climbs by roughly the size of a preview image per asset and never comes back down.
  5. Same with Jobs → OCR → All/Missing.

Relevant log output

Kernel, at the moment of the kill:

Out of memory: Killed process 3637 (immich) total-vm:34393088kB, anon-rss:6300768kB, file-rss:0kB, shmem-rss:0kB, UID:0 pgtables:28272kB oom_score_adj:0

Immich logs showed nothing unusual before that, other than the earlier "Successfully updated database CLIP dimension size from 512 to 768" when the model was changed. Once memory ran out, DNS lookups inside the containers started timing out and Docker health checks failed, but that is a symptom, not the cause.

Additional information

Reading the server source for v3.2.0, the one piece of code shared by Smart Search and OCR (and not by facial recognition, which I did not run) is MachineLearningRepository.predict() in server/src/repositories/machine-learning.repository.ts. Two things there look relevant:

  1. getFormData() (around line 233) does readFile(path) and then new Blob([new Uint8Array(fileBuffer)]). That makes three copies of each preview: the Buffer from readFile, a full copy from new Uint8Array(buffer) (the constructor copies a typed array, it does not wrap it), and a third copy into the Blob's native storage. As far as I can tell, memory held by Blobs is not counted by V8 as external memory, so the GC sees a small heap, rarely runs a full collection, and the native copies accumulate. That would explain per-asset growth that does not depend on concurrency and is not released between jobs.
  2. In predict() (around line 171), when the ML server returns a non-OK status the response body is never read or cancelled, and the fetch has no abort signal (the health check at line 133 does use one). Node/undici keeps unconsumed bodies alive, so this would make things worse whenever the ML server errors, though it is not the baseline growth.

Possibly related: #27617 (server RSS explodes and gets OOM-killed, closed as duplicate of #10023). #23462 is the Python ML container's leak and looks unrelated.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions