Skip to content

Unclear error and no recovery path when antivirus quarantines a bundled executable #13011

Description

@kelsos

Unclear error and no recovery path when antivirus quarantines a bundled executable

What happened

On Windows, a Defender definition set flagged the packaged colibri.exe as malware, under a
machine-learning heuristic detection rather than a signature match. This was a false positive.
rotki was running at the time, so Defender terminated the live colibri process and quarantined
the file from the installation directory underneath the running app. Defender's report named both
the executable and the running process id, so both the process and the file were taken in the
same action.

Microsoft removed the detection two definition revisions later, so there is nothing to work
around for this specific case. The problem worth fixing is that rotki handled the situation
badly, and it will recur: this applies to any security product, and a machine-learning heuristic
can flag any future build. Unsigned native binaries are routine false-positive material.

The problem

The incident produces two distinct failure states, and both surface to the user as an opaque
message that gives them no way to understand what happened or what to do about it.

1. colibri killed while rotki is running

colibri uses the default restart policy, max_retries: 0 / on_crash: ExitSupervisor
(crates/starling-core/src/config.rs:97-111), so its death tears down the whole supervisor. The
user is told only that the rotki backend stopped unexpectedly and that they should check the logs
(frontend/app/electron/main/starling-handler.ts:244-249).

2. rotki restarted afterwards, with the binary still missing

cmd.spawn() (crates/starling-core/src/process.rs:182) returns ENOENT. It propagates as
SupervisorError::Io, whose Display impl (crates/starling-core/src/error.rs:15) renders only
the bare io error, with no service name and no path. The user gets the full-screen
StartupErrorScreen, headed "rotki failed to start" / "There is a problem with the backend",
whose body is a raw "failed to start the backend: io error: ... (os error 2)" string and whose
only control is a "Terminate" button.

Because nothing in that message identifies which file is missing, it is indistinguishable from
a missing core or mcp binary, and nothing hints at a cause. The screen also instructs the user
to open a GitHub issue and attach logs, for something that is not a rotki bug and that they could
resolve themselves in a couple of minutes.

A file that is present but not executable (EACCES, e.g. blocked rather than deleted) takes the
same path and is equally opaque.

This is not specific to colibri

colibri is simply the binary that happened to be flagged. We ship several native executables, and
they do not all fail through the same code path, so a fix aimed only at colibri would leave gaps.

rotki-core (PyInstaller): spawned by the supervisor, same as colibri. This is arguably the
more likely candidate for a future false positive: one-file PyInstaller executables unpack a
bundled interpreter to a temp directory at startup, which is a textbook heuristic trigger, and
they are flagged more often than Rust binaries. Note the missing-core case behaves differently
from the missing-colibri case today: resolveCoreBinary()
(frontend/app/shared/starling/starling-paths.ts:35-42) does check and throw, but it throws
inside buildStarlingInvocation(), which is not wrapped in a try/catch at
frontend/app/electron/main/starling-handler.ts:160. So that one propagates as an uncaught
exception rather than reaching the error screen at all.

starling itself: this is the important asymmetry. starling is spawned by Electron directly
(frontend/app/shared/starling/starling-launch.ts:43-76), not by the supervisor. If starling is
the binary that gets quarantined, there is no supervisor left to report anything, and none of the
supervisor-side improvements apply. That check has to live in the Electron main process; nothing
else can cover it.

mcp: same supervisor path as colibri, and it already has a softer restart policy
(ReportOnly, crates/starling-core/src/config.rs:517-522), so it degrades rather than taking
the app down. Worth confirming its degraded state is legible to the user, and worth confirming
whether it ships as a separate executable or is bundled into another one.

The messaging and the recovery docs should therefore be written in terms of "a required rotki
component was removed", parameterised by which one, rather than hardcoded to colibri.

What the logs show

Two details from the incident logs matter for the fix.

The service name already exists and is discarded. The supervisor logs that colibri
specifically exited unexpectedly, naming the service. The message that reaches the user is the
generic "backend stopped unexpectedly". The same applies on the start path, where the resolved
colibri binary path is present in the spawn line that is already logged. Most of the messaging fix
is propagating information we already have, not detecting anything new.

The antivirus kill produced exit code 0. The terminated process was reported as crashing with
a zero exit code, which the main process then surfaced as a crash with a code that reads as if
nothing went wrong. A non-zero exit code therefore cannot be used to distinguish "killed by
antivirus" from a clean shutdown. The discriminator has to be a binary-existence check.

There is currently no mention of antivirus, quarantine, or false positives anywhere in the
repository, in code or in documentation.

What we want

1. Clear messaging

When a service binary is missing or not executable, say so plainly instead of surfacing a raw io
error:

  • Name the component (colibri, rotki-core, starling, mcp) and the exact path that was
    expected.
  • State whether the file is missing, or present but not executable.
  • Work for all four components, including starling itself, whose check has to live in Electron
    main because a missing starling means there is no supervisor to report anything.
  • On Windows, state directly that antivirus quarantine is the most likely cause, and that this is
    a known false-positive pattern rather than a sign the machine is compromised.
  • Cover both failure states above, so the user gets the same explanation whether the file vanished
    mid-session or rotki was restarted afterwards.
  • Do not tell the user to file a GitHub issue in this case. Point them at the recovery steps.

2. A recovery path

Link from the error screen to a documented troubleshooting section.

Ordering matters: restoring a file from quarantine while the offending definition set is still
active just gets it quarantined again. The generic steps are:

  1. Update antivirus definitions first.
  2. Restore the file from quarantine.
  3. If it recurs, add an exclusion for the rotki installation directory.
  4. If the file cannot be restored, reinstall rotki.

3. Documentation

Add a troubleshooting page covering antivirus quarantine. Keep the body vendor-neutral, since any
security product can do this, but include concrete Windows Defender instructions as the default
shipped scanner on Windows.

The recovery procedure does not need to vary by component. All the bundled executables live
under a single installation directory, so one folder exclusion covers every binary we ship, and
"restore from quarantine" is the same gesture whichever file was taken. Write one procedure, not
four. The component name belongs in the error message (so the user can find the right entry in
Protection history), not in a branching set of doc sections.

Suggested Defender-specific content:

Update definitions
Windows Security → Virus & threat protection → Protection updates → Check for updates.

Restore the quarantined file
Windows Security → Virus & threat protection → Protection history → locate the quarantined item
naming the file from the rotki error message → Actions → Restore.

Note that Defender removes quarantined files automatically after a retention period, so if enough
time has passed the file will be gone and reinstalling rotki is the remedy.

Add an exclusion (only if the detection recurs)
Windows Security → Virus & threat protection → Virus & threat protection settings → Manage
settings → Exclusions → Add or remove exclusions → Add an exclusion → Folder, then select the
rotki installation directory, which by default lives under %LOCALAPPDATA%\Programs\rotki.

The docs should state plainly that an exclusion reduces protection for that folder and should only
be used when the user is confident the detection is a false positive, e.g. after confirming the
release was downloaded from the official source.

Reporting the false positive
A short note that users can submit the file to Microsoft at
https://www.microsoft.com/en-us/wdsi/filesubmission to help get the detection removed, and that
reporting it to rotki lets us do the same.

Suggested implementation

The plumbing for this already exists; this is mostly wiring rather than new infrastructure.

  • Add existence/executability checks for every packaged binary, mirroring the existing
    resolveCoreBinary() at frontend/app/shared/starling/starling-paths.ts:35-42. The packaged
    colibri path is currently passed through unchecked at
    frontend/app/shared/starling/starling-launchers.ts:202-216, and the existing core check throws
    somewhere nothing catches it (see above). A single helper returning a structured
    "missing / not executable / ok" result per component, rather than throwing, would let all of
    these report through the same screen.
  • Check starling itself before spawning it in Electron main
    (frontend/app/shared/starling/starling-launch.ts:43-76). This is the one case no
    supervisor-side change can cover.
  • Enrich the supervisor-side spawn error to carry the service name and binary path
    (crates/starling-core/src/lifecycle.rs:283, crates/starling-core/src/error.rs:15-16), so an
    ENOENT from any service identifies itself. Note lifecycle.rs:283 also leaves the service in
    Spawning without setting last_error, unlike the readiness-failure branch at
    lifecycle.rs:298-303.
  • On the crash path (crates/starling-core/src/control/controller.rs:598-605 →
    starling-handler.ts:244-249), carry the service name through to the user-facing message (the
    controller already has it), and check whether the binary still exists before reporting a generic
    unexpected stop. That existence check is what distinguishes "colibri crashed" from "colibri was
    taken by antivirus". It cannot be done by exit code, as noted above.
  • Add a BackendCode.MISSING_BINARY (frontend/app/shared/ipc.ts:5-15) and a dedicated screen
    alongside MacOsVersionUnsupported.vue / WinVersionUnsupported.vue. It plugs into the existing
    handler at frontend/app/src/modules/shell/app/use-backend-messages.ts:62-75. There is precedent
    for a long, actionable, human-written message on this channel: the dev-proxy port conflict at
    starling-handler.ts:140-148.

Notes

  • Code signing would reduce recurrence but does not help a user whose file has already been
    quarantined, so it is a separate concern from this issue.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

WindowsbugSomething isn't workingdocumentationMissing or incorrect documentation that requires updatingfrontendrustrust related work

Projects

No projects

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions