Skip to content

RFC: should an unhealthy worker replica restart without an operator? #361

Description

@HMarzban

Parent

#328. Related: #154, #157

What to build

The hocuspocus-worker service runs queued jobs. Its store jobs write people's edits to the database. When one worker replica stops taking jobs, edits reach the database late until someone restarts it. Docker marks the container unhealthy but keeps it running. This RFC asks whether an unhealthy worker restarts without an operator, and which signal may trigger that.

Acceptance criteria

  • Alert when one worker replica stops dequeuing while jobs wait #154 is closed, so a stalled replica is visible on its own.
  • The maintainer picks one option below in a comment on this issue, and names the signal that triggers the restart.
  • apps/hocuspocus.server/CLAUDE.md §Production And Docker Compose records the ruling.
  • A build issue exists for the chosen option, or the ruling says the restart stays manual.

Blocked by

Agent brief

Type: HITL — the maintainer picks an option and its trigger signal, and writes it as a comment on this issue. An agent then records it in apps/hocuspocus.server/CLAUDE.md.

Category: ruling

Current behavior:

  • The hocuspocus-worker service runs two replicas with restart: always (docker-compose.prod.yml:389-477). Its healthcheck calls GET /health on port 4002 every 15 s, with 3 retries.
  • Docker restarts a container only when its process exits. A failing healthcheck marks it unhealthy and leaves it running. No compose file runs a service that restarts unhealthy containers.
  • /health returns unhealthy when one check fails (apps/hocuspocus.server/src/hocuspocus.worker.ts:154-223). The checks are:
    • the document worker runs and is not paused;
    • the oldest waiting store job is under 120 s;
    • the push and email consumers run;
    • Postgres and Redis answer.
  • The oldest-waiting check reads the shared wait list (getStoreQueueOldestWaitingAgeMs, apps/hocuspocus.server/src/lib/queue.ts:369). Both replicas read the same value. A stalled replica beside a draining sibling still reports healthy.
  • A Postgres or Redis outage makes both replicas unhealthy at once. A restart does not fix that, and it cuts jobs in flight.
  • The comment at hocuspocus.worker.ts:148-151 records the 2026-07-14 outage: BullMQ flags stayed green while the fetch loop was parked.

Desired behavior: A worker replica that stops taking jobs comes back without an operator. Or the ruling says the restart stays manual, and the alert is enough.

Options.

  • A. Keep it manual. The Alert when one worker replica stops dequeuing while jobs wait #154 alert pages the operator, who restarts the replica.
  • B. Self-exit on a local signal. The worker exits after N failed checks in a row of signals that belong to this process only. restart: always then brings it back. Shared signals, such as Postgres, Redis and the shared wait list, never trigger an exit.
  • C. A restart service. Run a container that restarts any container marked unhealthy. It needs the Docker socket mounted into that container, which the maintainer must accept. It also restarts both replicas during a database outage, unless the health check first drops the shared signals.

Evidence the maintainer needs.

  1. How often a worker replica has stalled in production, and how long each stall lasted.
  2. How often Alert when one worker replica stops dequeuing while jobs wait #154's alert fires once it is live.
  3. Which per-process signal exists for option B. The worker exposes no per-replica "last job taken" time today.

Where to start: The hocuspocus-worker service in docker-compose.prod.yml; healthApp.get('/health', …) in apps/hocuspocus.server/src/hocuspocus.worker.ts; getStoreQueueOldestWaitingAgeMs in apps/hocuspocus.server/src/lib/queue.ts. Line numbers are hints as of 2026-09-28; the agent searches by symbol.

Rules that apply:

  • apps/hocuspocus.server/CLAUDE.md §Production And Docker Compose.
  • The "Rolling-Deploy Strategy" header in docker-compose.prod.yml (do not change the deploy order).
  • .cursor/skills/tech-writer/SKILL.md §Simplified English (house standard) for the doc edit.

Verify: After the doc edit, run bun run check:agent-docs, then bun run check, from the repo root. Both pass. The ruling bullet appears under §Production And Docker Compose and links this issue.

Out of scope

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions