Skip to content

Silent data loss: external data drive unmount/remount → stop/start re-initialized a fresh empty data.img.raw (2.2.3) #2723

Description

@jrcheong

Related issues: #1726, #2090, #2364, #2450, #2189, #2628

Environment

  • OrbStack 2.2.3 (2020300), commit c83556b0ef8f1ba9a33abbb194622b6b7a1c0307
  • macOS 27.0 (26A428), Mac Studio M4 Max
  • Data directory on an external APFS SSD: /Volumes/ExternalSSD/orbstack (set in ~/.orbstack/vmconfig.json, data_dir unchanged throughout)
  • 15 containers + all volumes lived in data.img.raw (≈1.8 TB sparse)

Timeline (Oct 1, local)

  1. ~18:35 — external SSD briefly unmounted/remounted while OrbStack was running. The VM's /data mount went away; the "Sync data" health check failed with ENOTTY (seen in vmgr.log at the time; that log generation has since been rotated out — the lines below are from the surviving rotated log).
  2. I ran orbctl stop, then orbctl start.
  3. On start, vmgr ran the fresh-init path — not the "already initialized" path — and grew the new 1 GB image to the configured size. All 15 containers and all volumes were gone. No dialog, no prompt, no notification.
  4. Data was recovered from a pre-migration Docker Desktop Docker.raw copy (Sep 28 state). The 3 days of writes under OrbStack in between are unrecoverable.

Log evidence

~/.orbstack/log/vmgr.1.log (the start at 18:54):

18:54:28 startup phase begin  phase=initialize_data_image
18:54:28 initializing data
18:54:28 data image initialization complete
18:54:28 startup phase begin  phase=lock_data_image
18:54:28 resized data image GPT image_from=1074807296 image_to=1995000791552 partition_from=1071644672 partition_to=1994999726080
kernel [0.594917] BTRFS info (device vdb1): first mount of filesystem a1bc685b-8e1e-4a53-be23-c26036fb5dba
kernel [0.830954] BTRFS info (device vdb1): resize device /dev/vdb1 (devid 1) from 1071644672 to 1994999726080

For contrast, the next start (19:37, vmgr.log) shows the normal path:

19:37:21 startup phase begin  phase=initialize_data_image
19:37:21 data image already initialized

Notes:

What I expected

Per #1726 / #2090 / #2364, when the data image can't be validated, OrbStack should refuse to start (or show the "empty data — delete everything and reset?" dialog) rather than silently re-initialize. v2.2.2's release notes mention "Better protection against data corruption, with automatic recovery" — I suspect that recovery path is what re-initialized the image without asking.

Questions

  1. Under what conditions does vmgr choose initializing data vs data image already initialized? What validity check failed here after an unmount/remount of the backing volume?
  2. Is silent re-initialization intended behavior of the v2.2.2 "automatic recovery", or a regression vs the OrbStack falsely claims that data is empty after using Migration Assistant #2364 dialog?
  3. Is there any supported way to make the data image immutable-on-invalid (fail closed), or to keep a backup copy before any automatic recovery runs?

Happy to provide more logs. This cost me 3 days of container/volume writes; the silent part is what stings.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions