| id | ops-deployment | ||||
|---|---|---|---|---|---|
| type | ops | ||||
| status | accepted | ||||
| owns |
|
||||
| read_when |
|
||||
| tokens | 875 | ||||
| supersedes |
Two targets: Docker Compose for a single machine, Kubernetes for everything else. Both are first-class — self-hosting is the product (I-4), so neither is a toy.
just dc-up # nine infrastructure containers plus six services
just dc-status # health of each
just dc-logs # follow everything
just dc-down # stop
just dc-reset # stop and destroy volumesInfrastructure: nats, redpanda, redpanda-console, pd, tikv, scylladb, minio,
minio-init, envoy. Services: auth, messaging, presence, media, call, notification.
just dc-rebuild <service> rebuilds one image; just dc-shell <service> opens a shell;
just dc-cqlsh, just dc-tikv-status and just dc-redpanda-health reach the stores
directly.
Every service reaches its dependencies by Compose DNS name — pd:2379, scylladb:9042,
nats://nats:4222, minio:9000, redpanda:9092. Inside a container localhost is that
container itself, so a localhost address is always wrong here.
Redpanda advertises two listeners: internal://redpanda:9092 for clients on the Compose
network, and external://localhost:19092 for clients on the host. Only 19092 is published, so
in-container clients must use the internal one. auth-service (producer) and
messaging-service (consumer) are the only services using Kafka; both are given
REDPANDA_BROKERS. Host-side tooling such as infra/scripts/init-redpanda-topics.sh
correctly uses localhost:19092.
notification-service and call-service read plain environment variables in main.rs
rather than going through guardyn_common::config. notification-service reads LISTEN_ADDR
and SCYLLA_HOSTS; setting only GUARDYN_PORT and GUARDYN_DATABASE__SCYLLADB_NODES leaves
it binding the wrong port and dialling the Kubernetes ScyllaDB FQDN. Compose now sets both
forms. Unifying this is tracked separately.
messaging-service used to carry GUARDYN_E2EE_ENABLED, GUARDYN_MLS_ENABLED and four
companions in Compose and both k8s overlays. They are gone, and nothing replaces them.
The server is a pure relay (ADR-0010): the ratchet and
MLS run on the clients, it stores encrypted_content byte-for-byte and holds no key material,
so there is nothing left for such a flag to select. Invariant I-2 forbids one existing at all.
Two details worth keeping straight when reading older deployment files or git history:
- The dev overlay set
GUARDYN_E2EE_ENABLED: "false"with a comment offering it as a debugging convenience. Encryption was never something an operator could turn off for convenience. - The prod overlay set it to
"true", which read as though production encryption depended on it. It did not. After PR-32a/32b the code stopped consulting these variables entirely, so for a period they were inert while still appearing authoritative — which is precisely why they had to be deleted rather than left as harmless.
rules-verify enforces their absence (E2EE-FLAG), so reintroducing one fails CI.
just kube-create # k3d cluster from infra/k3d-config.yaml
just kube-bootstrap # cert-manager and core components
just k8s-deploy <service>
just verify-kube # smoke checks
just teardownThe local cluster is k3d: 3 servers, 2 agents, k3s v1.31.5-k3s1, Traefik disabled
because Envoy is the ingress path.
infra/k8s/base/ holds namespaces, apps, envoy, tikv, scylladb, minio,
cert-manager, cilium, monitoring and observability.
overlays/local is the development layer. overlays/prod adds hpa.yaml, pdb.yaml,
ingress.yaml, network-policies.yaml, service-monitors.yaml, slo-rules.yaml,
alertmanager-config.yaml and Grafana SLO dashboards.
| Service | Port |
|---|---|
| auth-service | 50051 |
| messaging-service | 50052 (gRPC), 8081 (WebSocket) |
| presence-service | 50053 |
| media-service | 50054 |
| call-service | 50056 (gRPC), 8085 |
| notification-service | 50055 |
Never hardcode a hostname. DOMAIN is the single source of truth, and every hostname
derives from it: auth.${DOMAIN}, api.${DOMAIN}, ws.${DOMAIN}, media.${DOMAIN},
app.${DOMAIN}. Deployment must work with .local, .test and real domains alike.
That rule has a second edge which is easy to miss: a URL in an annotation is still a
hardcoded domain. The runbook_url fields in infra/k8s/base/monitoring/alerting-rules.yaml
and infra/k8s/overlays/prod/slo-rules.yaml pointed at a project-owned host that served no
such path, so every one of them was a 404 waiting for an on-call engineer at 3am. They now
point at anchors in RUNBOOK.md in this repository, which is both domain-free
and the place the procedure actually lives. An operator's runbook link should resolve without
depending on who owns which domain this year.
SOPS with age, configured in .sops.yaml. Only *.enc.yaml is committed; age-key.txt
and any *.key are gitignored and must never reach the repository.
| Gap | Consequence | Owned by |
|---|---|---|
No call-service Deployment in infra/k8s/base/apps/ |
call-service is Compose-only and cannot be deployed to Kubernetes | PR-44 |
| Envoy routes 3 of 6 services (auth, messaging, presence) | media, calls and notifications are unreachable from a browser client | PR-44 |
infra/k8s/base/envoy/ingress.yaml:18 hardcodes envoy.guardyn.local |
breaks the ${DOMAIN} rule above |
unowned |
| Production images are tagged, not digest-pinned | a tag can be moved under a running cluster | PR-44 |
infra/secrets/.gitignore ignores *.enc.yaml — the encrypted file — while the plaintext app-secrets.yaml is tracked |
exactly inverted: the safe artefact is excluded and the unsafe one committed. The tracked values are placeholders, so no live credential is exposed yet | PR-42 |
infra/justfile is a second, divergent task file whose k8s:deploy references a values file that does not exist |
dead code that will mislead | unowned |