You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
The hocuspocus-worker service runs queued jobs. Its store jobs write people's edits to the database. When one worker replica stops taking jobs, edits reach the database late until someone restarts it. Docker marks the container unhealthy but keeps it running. This RFC asks whether an unhealthy worker restarts without an operator, and which signal may trigger that.
Type: HITL — the maintainer picks an option and its trigger signal, and writes it as a comment on this issue. An agent then records it in apps/hocuspocus.server/CLAUDE.md.
Category: ruling
Current behavior:
The hocuspocus-worker service runs two replicas with restart: always (docker-compose.prod.yml:389-477). Its healthcheck calls GET /health on port 4002 every 15 s, with 3 retries.
Docker restarts a container only when its process exits. A failing healthcheck marks it unhealthy and leaves it running. No compose file runs a service that restarts unhealthy containers.
/health returns unhealthy when one check fails (apps/hocuspocus.server/src/hocuspocus.worker.ts:154-223). The checks are:
the document worker runs and is not paused;
the oldest waiting store job is under 120 s;
the push and email consumers run;
Postgres and Redis answer.
The oldest-waiting check reads the shared wait list (getStoreQueueOldestWaitingAgeMs, apps/hocuspocus.server/src/lib/queue.ts:369). Both replicas read the same value. A stalled replica beside a draining sibling still reports healthy.
A Postgres or Redis outage makes both replicas unhealthy at once. A restart does not fix that, and it cuts jobs in flight.
The comment at hocuspocus.worker.ts:148-151 records the 2026-07-14 outage: BullMQ flags stayed green while the fetch loop was parked.
Desired behavior: A worker replica that stops taking jobs comes back without an operator. Or the ruling says the restart stays manual, and the alert is enough.
B. Self-exit on a local signal. The worker exits after N failed checks in a row of signals that belong to this process only. restart: always then brings it back. Shared signals, such as Postgres, Redis and the shared wait list, never trigger an exit.
C. A restart service. Run a container that restarts any container marked unhealthy. It needs the Docker socket mounted into that container, which the maintainer must accept. It also restarts both replicas during a database outage, unless the health check first drops the shared signals.
Evidence the maintainer needs.
How often a worker replica has stalled in production, and how long each stall lasted.
Which per-process signal exists for option B. The worker exposes no per-replica "last job taken" time today.
Where to start: The hocuspocus-worker service in docker-compose.prod.yml; healthApp.get('/health', …) in apps/hocuspocus.server/src/hocuspocus.worker.ts; getStoreQueueOldestWaitingAgeMs in apps/hocuspocus.server/src/lib/queue.ts. Line numbers are hints as of 2026-09-28; the agent searches by symbol.
Rules that apply:
apps/hocuspocus.server/CLAUDE.md §Production And Docker Compose.
The "Rolling-Deploy Strategy" header in docker-compose.prod.yml (do not change the deploy order).
.cursor/skills/tech-writer/SKILL.md §Simplified English (house standard) for the doc edit.
Verify: After the doc edit, run bun run check:agent-docs, then bun run check, from the repo root. Both pass. The ruling bullet appears under §Production And Docker Compose and links this issue.
Parent
#328. Related: #154, #157
What to build
The
hocuspocus-workerservice runs queued jobs. Its store jobs write people's edits to the database. When one worker replica stops taking jobs, edits reach the database late until someone restarts it. Docker marks the container unhealthy but keeps it running. This RFC asks whether an unhealthy worker restarts without an operator, and which signal may trigger that.Acceptance criteria
apps/hocuspocus.server/CLAUDE.md§Production And Docker Compose records the ruling.Blocked by
Agent brief
Type: HITL — the maintainer picks an option and its trigger signal, and writes it as a comment on this issue. An agent then records it in
apps/hocuspocus.server/CLAUDE.md.Category: ruling
Current behavior:
hocuspocus-workerservice runs two replicas withrestart: always(docker-compose.prod.yml:389-477). Its healthcheck callsGET /healthon port 4002 every 15 s, with 3 retries.unhealthyand leaves it running. No compose file runs a service that restarts unhealthy containers./healthreturnsunhealthywhen one check fails (apps/hocuspocus.server/src/hocuspocus.worker.ts:154-223). The checks are:getStoreQueueOldestWaitingAgeMs,apps/hocuspocus.server/src/lib/queue.ts:369). Both replicas read the same value. A stalled replica beside a draining sibling still reports healthy.hocuspocus.worker.ts:148-151records the 2026-07-14 outage: BullMQ flags stayed green while the fetch loop was parked.Desired behavior: A worker replica that stops taking jobs comes back without an operator. Or the ruling says the restart stays manual, and the alert is enough.
Options.
restart: alwaysthen brings it back. Shared signals, such as Postgres, Redis and the shared wait list, never trigger an exit.unhealthy. It needs the Docker socket mounted into that container, which the maintainer must accept. It also restarts both replicas during a database outage, unless the health check first drops the shared signals.Evidence the maintainer needs.
Where to start: The
hocuspocus-workerservice indocker-compose.prod.yml;healthApp.get('/health', …)inapps/hocuspocus.server/src/hocuspocus.worker.ts;getStoreQueueOldestWaitingAgeMsinapps/hocuspocus.server/src/lib/queue.ts. Line numbers are hints as of 2026-09-28; the agent searches by symbol.Rules that apply:
apps/hocuspocus.server/CLAUDE.md§Production And Docker Compose.docker-compose.prod.yml(do not change the deploy order)..cursor/skills/tech-writer/SKILL.md§Simplified English (house standard) for the doc edit.Verify: After the doc edit, run
bun run check:agent-docs, thenbun run check, from the repo root. Both pass. The ruling bullet appears under §Production And Docker Compose and links this issue.Out of scope
/health/readya caller. Worker /health/ready has no caller — give it one or delete it #157 owns it.