A self-hosted, open-source alternative to Seq for log aggregation and analysis. Drop-in compatible with Serilog.Sinks.Seq — point your existing sink at Log Cannon and keep shipping CLEF.
- Seq-compatible ingestion — works with existing Serilog/CLEF configurations, plus webhook and OpenTelemetry (logs + traces) endpoints.
- Edge ingestion that survives outages — logs are accepted at the Cloudflare edge and buffered in a queue, so nothing is dropped while your server is offline.
- ClickHouse storage — columnar storage tuned for high-volume time-series log data.
- Web dashboard — search, filter, and explore logs; build custom widget dashboards; manage API keys.
- Threshold alerts — define SQL conditions and get email when they trip.
- Per-service retention, backups, MCP — retention windows per source, twice-daily backups with offsite R2 sync, and an MCP server for AI assistants.
Log Cannon is a small monorepo of services. Ingestion rides the Cloudflare edge; everything else runs on your own server via Docker Compose.
Serilog / OTel / Webhooks
│
▼
┌─ Cloudflare edge ─────────────────────────┐
│ Ingest Worker ───────────► CF Queue │ validates API key (D1),
│ (CLEF · webhook · OTel) (buffers ≤4d) │ enqueues the raw payload
└──────────────────────────┬────────────────┘
│ pull
▼
┌─ Your server (Docker Compose) ────────────┐
│ Queue Consumer (Go) ──► ClickHouse │ parses CLEF/webhook/OTel,
│ Dashboard (Next.js) ──► reads ClickHouse │ batch-inserts
│ Alert Worker (Go) ────► email │
│ Retention Worker (Go) ─► trims ClickHouse │
│ Supabase Pull (Go) ───► ClickHouse │ optional: Supabase
│ Backup ───────────────► Cloudflare R2 │
│ │
│ Access via Cloudflare Tunnel (TLS/DDoS) │
└────────────────────────────────────────────┘
The Worker is deliberately thin — it only authenticates the request against its D1 key registry and pushes the raw body to the queue. All parsing happens in the Go queue consumer on your server, using the same code paths regardless of source format. When your server goes offline the Worker keeps accepting logs and the queue holds them (up to 4 days) until the consumer drains the backlog.
Cloudflare is a hard dependency. Ingestion requires a Cloudflare account with Workers and Queues; access uses a Cloudflare Tunnel. Storage, the dashboard, alerting, retention, and backups are fully self-hosted.
- Docker & Docker Compose
- A Cloudflare account with a domain (Workers free tier is sufficient)
- Wrangler CLI and pnpm to deploy the ingest Worker
- An SMTP provider (or HTTP email API) for alert and sign-in emails — optional for local dev, where a bundled mailbox catches everything
git clone <repo-url>
cd log-cannon
cp .env.example .env
# Edit .env — at minimum set AUTH_SECRET and AUTH_ALLOWED_EMAILS,
# plus CLOUDFLARE_TUNNEL_TOKEN and your email settings (see below).
docker compose up -dThis brings up ClickHouse, the dashboard, the alert/retention/backup workers, and the queue consumer. For local development, add COMPOSE_PROFILES=dev to also start a bundled Inbucket mailbox (read captured emails at http://localhost:9000).
The Worker is what your log clients actually talk to. See Edge Ingestion Setup below for the full walkthrough (create a Queue + D1 database, bootstrap an admin key, deploy).
Log.Logger = new LoggerConfiguration()
.WriteTo.Seq(
serverUrl: "https://logs.yourdomain.com",
apiKey: "your-api-key-here")
.CreateLogger();Or in appsettings.json:
{
"Serilog": {
"WriteTo": [{
"Name": "Seq",
"Args": {
"serverUrl": "https://logs.yourdomain.com",
"apiKey": "your-api-key-here"
}
}]
}
}The dashboard and any direct-to-server endpoints are exposed via a Cloudflare Tunnel, which provides TLS and DDoS protection without opening inbound ports.
- Cloudflare Zero Trust → Networks → Tunnels → Create a tunnel (name it
log-cannon). - Copy the token into
.envasCLOUDFLARE_TUNNEL_TOKEN. - Add a public hostname:
logs-dashboard.yourdomain.com→http://dashboard:3000.
The ingest hostname (e.g. logs.yourdomain.com) is served by the Cloudflare Worker, configured separately in Edge Ingestion Setup.
The dashboard uses HMAC-signed session cookies and 6-digit email OTPs — there are no user accounts to manage. Anyone whose address is in AUTH_ALLOWED_EMAILS can request a code and sign in. OTP records live in SQLite at /app/data/auth.db inside the dashboard container (a Docker volume persists them across restarts).
Set EMAIL_TRANSPORT to pick how codes are delivered:
EMAIL_TRANSPORT |
When to use | Needs |
|---|---|---|
smtp (default) |
Local dev (bundled Inbucket) or any SMTP provider (Resend, Mailgun, SES, Postmark…) | SMTP_HOST, SMTP_PORT |
saasmail |
A simple HTTP email API that accepts a multipart payload field |
SAASMAIL_API_KEY, SAASMAIL_API_URL |
For most deployments, smtp pointed at your provider is the simplest path.
Two things send email: dashboard sign-in OTPs and alert notifications.
- OTP emails honor
EMAIL_TRANSPORT(smtporsaasmail, above). - Alert emails are delivered via the HTTP email API (
SAASMAIL_API_KEY/SAASMAIL_API_URL). The endpoint receives aPOST {SAASMAIL_API_URL}/api/sendwith a multipartpayloadfield containing{to, fromAddress, subject, bodyText, bodyHtml}and aBearertoken. PointSAASMAIL_API_URLat any service that speaks this shape.
API keys live in D1, behind the ingest Worker — there is no separate store to keep in sync. Create and manage them from the API Keys page in the dashboard (it talks to the Worker's admin API), or directly against the Worker:
curl -X POST https://logs.yourdomain.com/v1/keys \
-H "X-Api-Key: your-admin-scoped-key" \
-H "Content-Type: application/json" \
-d '{"name":"my-app","scopes":"ingest"}'See Edge Ingestion Setup for bootstrapping the first admin key.
Changes: GET/POST /api/v1/keys (the dashboard's REST API) now return created_at as an ISO 8601 timestamp, not the previous ClickHouse-formatted string. The field name is unchanged.
Each API key has a retention_days setting (0 = keep forever, the default). logs.api_keys is no longer read by anything — do not hand-edit it; the dashboard's periodic projection sync will overwrite or ALTER ... DELETE anything there that doesn't match the D1 registry. Set retention through one of the two paths that actually work:
- The API Keys page in the dashboard, or
PATCH /v1/keys/{keyId}against the ingest Worker with an admin-scoped key:
curl -X PATCH https://logs.yourdomain.com/v1/keys/your-key-id \
-H "X-Api-Key: your-admin-scoped-key" \
-H "Content-Type: application/json" \
-d '{"retentionDays": 14}'Either path updates D1 immediately; the dashboard projects that change into ClickHouse logs.key_policies after every mutation and on a periodic sync (see dashboard/src/instrumentation.ts). The retention-worker trims logs older than the configured window once per RETENTION_INTERVAL_HOURS (default 24h), per source — so a retention change can take up to a day to actually trim anything, even though it reaches D1 and the projection right away. "I set 7 days and nothing happened yet" is expected until the worker's next pass.
Build dashboards from configurable widgets backed by raw SQL against ClickHouse.
| Type | Description | Key config |
|---|---|---|
stat |
Single KPI metric | valueField, format (number/percent/duration), trend |
line_chart |
Time-series line chart | xField, yField (string or array), colors |
bar_chart |
Categorical bar chart | xField, yField (string or array), colors |
pie_chart / doughnut_chart |
Proportional chart | xField (label), yField (value), colors |
scatter_chart |
Correlation scatter | xField (numeric), yField (numeric), colors |
table |
Sortable data table | columns, sortBy |
{
"layout": "auto",
"widgets": [
{
"id": "errors-by-source",
"type": "pie_chart",
"title": "Errors by Source",
"dataSource": {
"sql": "SELECT source as name, count() as count FROM logs.events WHERE level = 'Error' AND timestamp > now() - INTERVAL 24 HOUR GROUP BY source"
},
"visualization": {
"xField": "name",
"yField": "count",
"colors": ["#FF4D2A", "#FF3366", "#36A2EB", "#FFCE56"]
}
},
{
"id": "log-volume",
"type": "line_chart",
"title": "Log Volume (24h)",
"dataSource": {
"sql": "SELECT toStartOfHour(timestamp) as time, count() as count FROM logs.events WHERE timestamp > now() - INTERVAL 24 HOUR GROUP BY time ORDER BY time"
},
"visualization": { "xField": "time", "yField": "count" }
}
]
}Define alerts in alert-worker/alerts.json. Each alert runs a SQL query on an interval and emails recipients when its condition is met:
count()overlogs.eventscan over-count. Ingest is at-least-once and nothing dedupes:idis generated at insert, so if the consumer inserts a batch and then fails to acknowledge it, the queue redelivers and those events are inserted a second time as distinct rows. The consumer retries the ack inside the message lease to make this rare (queue-consumer/main.go), but it is not impossible. Prefer a threshold with margin over one that trips on an exact count, and treat a single doubled interval as suspect before treating it as real.
{
"id": "high-error-rate",
"name": "High Error Rate",
"query": "SELECT count() as cnt FROM logs.events WHERE level = 'Error' AND timestamp > now() - INTERVAL 5 MINUTE",
"condition": "cnt > 50",
"interval_seconds": 60,
"cooldown_seconds": 300,
"recipients": ["alerts@example.com"],
"subject": "High error rate detected"
}The dashboard exposes a REST API and an MCP server, both authenticated with the same API keys (scoped read or write).
| Endpoint | Method | Description |
|---|---|---|
/api/v1/* |
Various | REST API (queries, keys, dashboards, alerts, backups…) |
/api/mcp |
POST | MCP server (Model Context Protocol) |
Ingestion endpoints (/ingest/clef, /api/events/raw, /ingest/webhook, /ingest/otlp/*, /v1/logs, /v1/traces) are served by the edge Worker — see Ingest Routes.
Log Cannon exposes its API as an MCP server at /api/mcp, so AI assistants and other MCP clients can use its tools directly.
Claude Code (~/.claude.json) / Cursor (.cursor/mcp.json):
{
"mcpServers": {
"log-cannon": {
"type": "http",
"url": "https://your-instance/api/mcp",
"headers": { "X-Api-Key": "your-api-key" }
}
}
}Tools are scoped to your key's permissions. See the MCP page in the dashboard for interactive setup.
Your log clients talk to a single Cloudflare Worker that validates API keys (against D1) and enqueues raw payloads onto a Cloudflare Queue. The Go queue-consumer (already running in Compose) drains the queue into ClickHouse.
npx wrangler queues create log-cannon-ingest # note the Queue IDnpx wrangler d1 create log-cannon-keys # note the database ID
npx wrangler d1 migrations apply log-cannon-keys --remoteKeys are managed in the dashboard UI, or directly through the Worker's admin API. Bootstrap the first admin key by hand:
npx wrangler d1 execute log-cannon-keys --remote --command \
"INSERT INTO api_keys (api_key, key_id, name, enabled, scopes, retention_days, created_at)
VALUES ('your-admin-key', lower(hex(randomblob(16))), 'admin', 1, 'admin', 0, datetime('now'))"Set that key as LOG_CANNON_ADMIN_KEY in your .env so the dashboard can
manage keys. There is no second store to keep in sync.
Update workers/packages/ingest/wrangler.toml with your D1 database ID and routes:
[[d1_databases]]
binding = "KEYS_DB"
database_name = "log-cannon-keys"
database_id = "YOUR_D1_DATABASE_ID" # ← replace
migrations_dir = "migrations"
routes = [
{ pattern = "logs.yourdomain.com/ingest/*", zone_name = "yourdomain.com" },
{ pattern = "logs.yourdomain.com/api/events/raw", zone_name = "yourdomain.com" },
{ pattern = "logs.yourdomain.com/v1/*", zone_name = "yourdomain.com" },
]cd workers
pnpm install
cd packages/ingest && pnpm wrangler deployAdd the Cloudflare credentials to your .env (the consumer needs an API token with Queues: Read):
CF_ACCOUNT_ID=your-cloudflare-account-id
CF_QUEUE_ID=your-queue-id-from-step-1
CF_API_TOKEN=your-api-token-with-queue-permissionsThen restart it:
docker compose up -d queue-consumercurl -X POST https://logs.yourdomain.com/ingest/clef \
-H "X-Seq-ApiKey: your-api-key" \
-d '{"@t":"2026-01-01T00:00:00Z","@mt":"Hello from edge","@l":"Information"}'
docker compose logs -f queue-consumerA single Worker handles every ingestion format via path-based routing:
| Path | Format | Notes |
|---|---|---|
/ingest/clef |
CLEF (NDJSON) | Primary Seq/Serilog endpoint |
/api/events/raw |
CLEF (NDJSON) | Legacy Seq compatibility |
/ingest/webhook |
Webhook JSON | Supports ?preset=cloudflare |
/ingest/otlp/logs |
OTel logs | Protobuf or JSON |
/ingest/otlp/traces |
OTel traces | Protobuf or JSON |
/v1/logs |
OTel logs | Standard OTel SDK path |
/v1/traces |
OTel traces | Standard OTel SDK path |
/health |
— | Returns {"status":"ok"} |
Senders bound their outbound request — Serilog's Seq sink and most hand-rolled clients both do — so a stalled ingest call is aborted client-side. Nothing is written and the response is never read, which leaves the sender's timeout as the only evidence the request happened. Two things make it visible from the ingest side instead.
Workers Logs. [observability] in wrangler.toml is on, with
invocation_logs, so every request records wall time, CPU time and outcome.
That is where ingest p50/p95/p99 comes from, and it retains the warning below.
Note it does not set logs.destinations: pointing a Logpush destination at a
Log Cannon ingest route is right for other Workers, but on this one it loops —
the destination delivers to /ingest/webhook on the very Worker whose
invocations it is reporting.
Slow-request events. Any ingest request that spends SLOW_REQUEST_MS or
more in the Worker is written as a Warning to SLOW_REQUEST_SOURCE, through
the same queue as every other event, so it lands in logs.events alongside
everything else. The emission happens after the response, and is throttled per
isolate so a degraded queue is not answered with more traffic.
| Worker var | Default | Description |
|---|---|---|
SLOW_REQUEST_MS |
1000 |
Report at or above this many milliseconds. 0 disables. |
SLOW_REQUEST_SOURCE |
log-cannon-ingest |
Source the events are written to. Empty leaves Workers Logs as the only channel. |
QUEUE_ACK_DEADLINE_MS |
0 |
Stop waiting for the queue ack after this long and finish the send in the background. 0 waits in full. |
The Worker stamps this source itself and needs no key for it, but retention is
projected from the key registry by name (see Per-Service Retention),
so a source with no key of the same name is kept forever. Volume is low — only
requests over the threshold, throttled per isolate — but if you want these
trimmed, register a key named SLOW_REQUEST_SOURCE and set its retentionDays.
Each event carries the breakdown, which is the point — a 5-second request is a different incident depending on where the time went:
| Property | Meaning |
|---|---|
TotalMs |
Time inside the Worker |
AuthMs |
D1 key lookup |
ReadMs |
Reading (and gzip-decoding) the request body |
EnqueueMs |
Wait on the INGEST_QUEUE send — the full send, or the deadline if one was hit |
EnqueueHandedOff |
Present, and true, only when the send outlived QUEUE_ACK_DEADLINE_MS and finished in the background |
KeyCacheHit |
Whether the key was already in this isolate's cache |
Route / Format / Status |
Which endpoint, which parser, what was returned |
IngestSource |
The source the request was writing to |
BodyBytes / QueueMessages |
Payload size and how many queue messages it became |
Colo / RayId |
Cloudflare edge location and Ray ID, for correlating with Workers Logs |
A cold isolate pays for a D1 round trip on every request, so KeyCacheHit: false with a large AuthMs is the expected shape for a source that sends a
handful of events a day — worth ruling in before looking further.
TotalMs starts when the Worker is invoked, so connection setup, TLS and
isolate cold start sit outside it. A sender that saw 5 s against a small
TotalMs is a result, not a broken measurement: it places the stall before
the Worker rather than inside it. Cloudflare freezes Date.now() between I/O
operations, so these spans measure I/O; a CPU-bound stall shows up in Workers
Logs' cpuTime instead.
To alert on degradation while it is happening, add an alert over these events:
{
"id": "ingest-latency",
"name": "Ingest Latency",
"query": "SELECT count() as cnt FROM logs.events WHERE source = 'log-cannon-ingest' AND timestamp > now() - INTERVAL 5 MINUTE",
"condition": "cnt > 0",
"interval_seconds": 60,
"cooldown_seconds": 900
}Because the events are throttled per isolate, treat cnt as "how many isolates
saw this", not as a count of slow requests. For the latter, read Workers Logs.
The INGEST_QUEUE send is a durable-acknowledgement round trip, so whatever
Cloudflare Queues takes to acknowledge is time the sender spends waiting for its
response. On a healthy, idle queue that tail has been measured in seconds, with
a worst case over a minute — and because senders bound their own request
(Serilog's Seq sink and most hand-rolled clients at around 5 s), they abort and
lose the batch while the platform itself is fine.
QUEUE_ACK_DEADLINE_MS bounds the caller's wait, not the send. Set it, and a
send that has not acked by the deadline is moved to the background: the response
goes out immediately and the send runs to completion behind it. Chunked bodies
are safe — the deadline covers the whole enqueue, so the batches that had not
gone out yet still go out.
It ships at 0, meaning off, and should be switched on deliberately. A handoff
means the Worker has answered 201 before the queue has durably accepted the
payload, so a send that then fails is a loss the client already believes
succeeded. That case is reported, never silent: an ingest-enqueue-failed-after-response
line in Workers Logs and a matching CLEF Error on SLOW_REQUEST_SOURCE,
carrying the route, the source, and how many messages were lost. Alert on it
separately from the latency alert above:
{
"id": "ingest-enqueue-loss",
"name": "Ingest Enqueue Loss",
"query": "SELECT count() as cnt FROM logs.events WHERE source = 'log-cannon-ingest' AND level = 'Error' AND timestamp > now() - INTERVAL 5 MINUTE",
"condition": "cnt > 0",
"interval_seconds": 60,
"cooldown_seconds": 900
}Choosing a value. Size it from your own measured EnqueueMs, not from a
round number. Two bounds fix it: it must sit far enough below your senders'
own fetch timeout that the response still reaches them, and far enough above
your queue's ordinary ack that sends which would have completed in time are
untouched. Query the slow-request events for the distribution — SELECT quantile(0.9)(JSONExtractFloat(properties, 'EnqueueMs')) FROM logs.events WHERE source = 'log-cannon-ingest' — and pick a deadline above the bulk of it. A
deadline below your p90 hands off routinely and converts an ordinary slow send
into an un-acked one, which is the trade running backwards.
queue-consumer, alert-worker, and retention-worker still write to stdout
(docker compose logs), and can also ship CLEF into Log Cannon the same way
any other client does: POST {LOG_CANNON_INGEST_URL}/ingest/clef with
X-Api-Key. Shipping is best-effort and never blocks the service's real work;
if the URL or key is unset, the service starts normally with shipping off.
Mint one ingest-scoped key per service. The key's registry name becomes
logs.events.source — use these names so operators know what to query and
what to attach retention to:
| Service | Env var for the key | Key name / source |
|---|---|---|
| queue-consumer | LOG_CANNON_QUEUE_CONSUMER_API_KEY |
log-cannon-queue-consumer |
| alert-worker | LOG_CANNON_ALERT_WORKER_API_KEY |
log-cannon-alert-worker |
| retention-worker | LOG_CANNON_RETENTION_WORKER_API_KEY |
log-cannon-retention-worker |
| supabase-pull | LOG_CANNON_SUPABASE_PULL_API_KEY |
supabase-pull |
All four share LOG_CANNON_INGEST_URL (same as the dashboard). Compose maps
each *_API_KEY into the container as LOG_CANNON_API_KEY.
Feedback loop. The consumer is the process that inserts whatever it ships.
Logging every poll cycle about inserting N events would enqueue more events to
insert. So the consumer does not tee stdout: it only ships a CLEF event when
a poll is slow or failed — Warning when total time crosses POLL_SLOW_MS
(default 1000), Error on pull/flush failure (with the same phase breakdown)
— and at most once per POLL_TELEMETRY_MIN_INTERVAL_MS (default 10000).
Stdout still prints every pull/insert with durations. Alert and retention tee
every log.Printf — their volume is low.
A source with no matching key in logs.key_policies is kept forever; register
keys with the names above (and set retentionDays) if you want these trimmed.
| Consumer var | Default | Description |
|---|---|---|
POLL_SLOW_MS |
1000 |
Ship a timing event when any phase or the total is at least this many milliseconds. 0 disables CLEF timings (stdout unchanged). |
POLL_TELEMETRY_MIN_INTERVAL_MS |
10000 |
Minimum gap between two queue-borne slow-poll events. |
ASYNC_ACK_DEPTH |
0 |
Hand the ack to a background worker instead of waiting for it. 0 keeps acking inline. See below. |
QUEUE_CONSUMER_WORKERS |
1 |
Independent poll loops. 1 is the single serial loop. Clamped to the ClickHouse pool (10). See below. |
A poll is pull → parse → insert → ack, in series, and the next pull cannot
start until the ack lands. The ack is about half of it — ~1.2s of a ~2.5s cycle
at rest, and ~2.4s of ~5s while Cloudflare Queues is degraded — against an
insert that costs ~12ms. So consumer throughput tracks Queues latency rather
than the work being done, and falls exactly when the queue is filling faster.
ASYNC_ACK_DEPTH > 0 starts a background ack worker with a queue that deep.
The poll hands off and returns; the ack completes behind it. That is safe
against early redelivery because a message stays leased for
visibility_timeout_ms (120s) whether or not its ack has arrived — the lease
bounds it, not the ack.
It ships at 0, and that default is the point. With it on, poll() reports
a clean cycle before the queue has acknowledged, so a failing ack surfaces
asynchronously — a log line and a CLEF Error on the consumer's source —
rather than as the poll's own error. The failure is still loud and still
alertable; it just arrives after the poll said it was fine. That is a real
change to how a duplicate-producing failure is noticed, so turn it on
deliberately.
Two behaviours worth knowing:
- A full queue acks inline. If the worker falls behind, the poll acks synchronously rather than dropping the job or letting the backlog grow. A dropped ack is a guaranteed duplicate, and an unbounded backlog would push the ack past the lease and duplicate anyway.
- The depth is bounded by arithmetic, not taste. A handed-off job can wait
behind the whole queue plus one in flight, so it completes at worst
(depth+1) x 15safter a flush that may already have used 60s of the 120s lease. Anything deeper starts acking after the lease has expired — by which point the message has been redelivered and the batch is inserted twice. The consumer clampsASYNC_ACK_DEPTHto what fits (currently 2) and logs when it does. - Shutdown drains. Pending acks finish on their own context, not the
cancelled one, because the events are already in ClickHouse — an abandoned ack
is not a lost write but a duplicated one on the next start. The grace is
derived from the same arithmetic (~50s), and a poll still acking inline
(synchronously, or because the ack queue was full) can need its own 15s
before the drain starts. The synchronous ack also uses its own context, so
SIGTERM does not abandon it either.
docker-compose.ymlsetsstop_grace_period: 75sfor this service; on Docker's 10s default, SIGKILL would land mid-drain and the code's patience would be fiction.
AckMs measures the handoff rather than the round trip once this is on, so
TotalMs stops including the ack — that is the improvement, but it does change
what the numbers mean. Events carry AckAsync: true so the two regimes are
distinguishable in logs.events; an event from the default path is
byte-identical to one from before this existed.
With the ack moved off the critical path, the pull is the remaining serial cost. A pull to Cloudflare Queues takes 1–2s at rest, and several times a day it fails with a 504 after ~15s, which freezes the only loop for that long.
QUEUE_CONSUMER_WORKERS=N runs N independent poll loops. Each one does the
same pull → parse → insert → ack-after-flush cycle on its own lease, so
Queues latency is paid in parallel and one stuck pull holds up 1/N of
intake instead of all of it.
- Ack ordering is unchanged. Pulls lease disjoint sets of messages for the 120s visibility window, so each worker keeps the ack-after-flush invariant for its own messages. No worker acks anything another worker inserted.
- Order is not a concern. Every event carries its own
timestamp;logs.eventssorts by it. - Shared ack worker. With
ASYNC_ACK_DEPTHon, all workers hand off to the same bounded ack queue, and a full queue still acks inline, so the lease arithmetic above holds for any N. - The cost to watch is ClickHouse inserts. N workers make N smaller inserts
per cycle.
InsertMsis in the slow-poll events; if it starts rising, stop raising N. N is clamped to the ClickHouse connection pool (10) so a flush never waits for a connection. - Idle cost. An empty queue gets up to N pulls a second instead of one. Pull consumers are limited per queue by message throughput (5,000/s), not by the general API request limit.
It ships at 1: this is the only writer of logs.events, and N workers means
N times the ack calls, so N times the exposure to a failed ack duplicating a
batch (#116). Raise it deliberately, and measure with the ingest-lag query in
#111.
Slow-poll properties: Messages, Events, PullMs, ParseMs, InsertMs,
AckMs, TotalMs.
A message the consumer cannot insert is retried, and if it still fails it is
parked on log-cannon-ingest-dlq rather than deleted. Without that queue,
Cloudflare discards a message once it exhausts max_retries — and for a log
platform the discarded payload is exactly the evidence you would want.
Current settings on the log-cannon-ingest pull consumer:
| Setting | Value | Why |
|---|---|---|
max_retries |
5 |
~10 minutes at the 120s lease — enough to ride out a ClickHouse restart or a Portainer redeploy without anyone touching the DLQ. |
visibility_timeout_ms |
120000 |
The lease. flushTimeout (60s) + the ack retry budget (15s) must fit inside it. |
batch_size |
100 |
Matches what each pull requests. |
dead_letter_queue |
log-cannon-ingest-dlq |
Where a message goes after max_retries. |
The DLQ keeps messages for 14 days (not the 4-day default), so something parked on a Friday is still there on Monday.
Two shapes, and they want different responses:
- One message, repeatedly. Almost certainly a payload the consumer cannot insert — a bad timestamp, an oversized row. Since an append failure is isolated to its own message, it parks alone and nothing else is affected. Inspect it, fix the cause, and decide whether to replay it.
- Many messages at once. Not a payload problem — the write path was down for longer than the retry budget. Check ClickHouse first, then replay.
The DLQ has an http_pull consumer attached purely so it can be read; nothing
polls it. Pull without acking to look without consuming — the messages return
after the visibility timeout:
ACCT=<account id>; DLQ=<dlq queue id>
curl -s -X POST \
"https://api.cloudflare.com/client/v4/accounts/$ACCT/queues/$DLQ/messages/pull" \
-H "Authorization: Bearer $CF_API_TOKEN" \
-H "Content-Type: application/json" \
-d '{"visibility_timeout_ms": 30000, "batch_size": 10}' | jq '.result.messages[].body'Each message body is the envelope the ingest Worker enqueues — source,
format, contentType, and the raw request body base64-encoded — so the
original payload is recoverable.
One wrinkle: the pull API sometimes returns that envelope double-encoded, as
a JSON string rather than an object (queue-consumer/main.go unwraps this on
every message). Handle both:
... | jq -r '.result.messages[0].body | if type == "string" then fromjson else . end | .body' | base64 -dThat prints the exact bytes the client POSTed — for CLEF, one JSON event per line.
There is no automatic replay, deliberately — a message is in the DLQ because
inserting it failed, and replaying blindly just fails again. Once the cause is
fixed, replay by POSTing the decoded payload back to the normal ingest endpoint
(/ingest/clef with an ingest key), then ack the DLQ copy so it does not linger.
Ack with the lease_id from the pull:
curl -s -X POST \
"https://api.cloudflare.com/client/v4/accounts/$ACCT/queues/$DLQ/messages/ack" \
-H "Authorization: Bearer $CF_API_TOKEN" \
-H "Content-Type: application/json" \
-d '{"acks": [{"lease_id": "<lease_id>"}]}'If the messages are not worth replaying, acking them is how you discard them.
supabase-pull pulls a Supabase project's platform logs (Postgres, Auth, Edge
Function console output and Edge Function requests) from the Management API's
unified logs endpoint and writes them straight into logs.events. It runs in
Compose next to the other workers and does nothing until a project is enabled.
SUPABASE_PULL_PROJECTS=abcdefghijklmnopqrst=myapp-supabase,<ref>=<source>
SUPABASE_ACCESS_TOKEN=sbp_... # personal access token with analytics_logs_readEach row lands with source = <source> and these properties, which is what
alerts and queries should filter on:
| Property | Value |
|---|---|
Source |
the configured source name (myapp-supabase) |
LogTable |
postgres_logs, auth_logs, function_logs or function_edge_logs |
ProjectRef |
the Supabase project ref |
sb_id |
the Supabase row id — the dedupe key |
sb_<attribute> |
every non-empty log_attributes key, dots and other non-word characters flattened to _ (request.path → sb_request_path) |
level maps from severity_text (ERROR → Error, LOG/NOTICE/INFO →
Information, PANIC/FATAL → Fatal, and so on); message is the
event_message.
How it keeps its place. There is no state file. On start it takes
max(timestamp) per LogTable from logs.events for rows whose Source
property matches, which lets it resume from rows another shipper wrote under a
different source column. Each run then re-reads from watermark − SUPABASE_PULL_OVERLAP so late-arriving rows are still collected, and skips any
row whose sb_id is already stored (rows without one, from an earlier shipper,
are matched on timestamp and message). Stopping it for a while and starting it
again fills the gap without duplicating rows, back as far as
SUPABASE_PULL_MAX_LOOKBACK. Keep that within the log retention of your
Supabase plan.
Catch-up. A run pages (500 rows, keyset on timestamp and id) until it is
caught up, walking 23-hour windows because the endpoint rejects a longer one,
and stops at SUPABASE_PULL_MAX_PAGES per table (default 50) so one huge
backlog cannot hold a run indefinitely. Where it stopped is where the next run
starts.
Rate limit. The logs endpoint allows 10 requests per window per project
and answers 429 with Retry-After beyond that. Steady state is one request
per table per run (4 a minute per project at the defaults), which fits. A
catch-up that needs more pages waits out Retry-After and retries (up to 5
times) instead of failing the table, so a long backlog drains at the API's
pace. Anything else calling the same project's logs endpoint shares that
budget.
Self-report. One CLEF event per run under the supabase-pull key:
Inserted, Pages, MaxReadLagSeconds (how far behind the present the worst table has been read — near zero when caught up), Behind, and a Tables array
with per-table rows, pages, read lag and watermark age (the newest stored row; absent for a table with none). Any failed table (API error,
refused token, ClickHouse error) raises the event to Error, so
source = 'supabase-pull' AND level = 'Error' is the alert condition for the
puller itself. Stdout carries one line per table per run.
Retention. Register an ingest key named after each <source> (and one named
supabase-pull) and set retentionDays. The puller doesn't use those keys to
write; retention is looked up by source name, and a source with no key is kept
forever.
Do not run two shippers for one project. Anything else that ships the same project's logs (Log Drains, a cron poller) must be stopped before the project is added here, or every event arrives twice. Deploy with the project list empty, stop the other shipper, then enable the project.
| Variable | Default | Description |
|---|---|---|
SUPABASE_PULL_PROJECTS |
(empty) | ref=source pairs, comma-separated. Empty = idle. |
SUPABASE_ACCESS_TOKEN |
— | Supabase PAT; required once a project is enabled |
SUPABASE_PULL_TABLES |
all four | Comma-separated subset of the tables above |
SUPABASE_PULL_INTERVAL |
1m |
Pause between runs |
SUPABASE_PULL_OVERLAP |
5m |
How far behind the watermark each run re-reads |
SUPABASE_PULL_MAX_LOOKBACK |
72h |
Oldest a run will reach, including after an outage |
SUPABASE_PULL_MAX_PAGES |
50 |
Pages per table per run |
Automated ClickHouse backups run twice daily with offsite sync to Cloudflare R2.
- Cloudflare dashboard → R2 Object Storage → Create bucket (e.g.
log-cannon-backups). - Manage R2 API Tokens → Create API Token with Object Read & Write on the bucket.
- Note the Account ID, Access Key ID, and Secret Access Key, and add them to
.env:
R2_ACCOUNT_ID=your-account-id
R2_ACCESS_KEY_ID=your-access-key-id
R2_SECRET_ACCESS_KEY=your-secret-access-key
R2_BUCKET=log-cannon-backupsThen rebuild: docker compose build backup && docker compose up -d backup.
| Variable | Default | Description |
|---|---|---|
BACKUP_CRON |
0 3,15 * * * |
Cron schedule (3 AM and 3 PM) |
BACKUP_RETAIN_LOCAL |
7 |
Local backups to keep |
BACKUP_RETAIN_OFFSITE |
14 |
R2 backups to keep |
R2_BUCKET |
log-cannon-backups |
R2 bucket name |
docker compose exec backup /scripts/backup.sh # manual backup
docker compose exec backup /scripts/restore.sh # list available backups
docker compose exec backup /scripts/restore.sh logs-full-2026-03-15-030000 # restore (auto-downloads from R2 if not local)Restore replaces, it does not append. restore.sh drops the logs database before restoring so a run against a live volume does not duplicate events. ClickHouse’s allow_non_empty_tables setting is append-only (it inserts into existing tables and can silently double history); the script uses it only when applying an incremental backup’s delta onto a freshly restored base. Do not run restore against production unless you intend a full replace.
Disaster recovery (fresh server): set up Docker Compose, clone the repo, configure .env with your R2 credentials, docker compose up -d, wait for ClickHouse to go healthy, then run the restore command above. Backup status is also visible in the dashboard under System → Backups.
See .env.example for the full annotated list. The essentials:
| Variable | Required | Description |
|---|---|---|
AUTH_SECRET |
Yes | 32+ random bytes signing session cookies (openssl rand -hex 32) |
AUTH_ALLOWED_EMAILS |
Yes | Comma-separated allowlist of sign-in emails |
CLOUDFLARE_TUNNEL_TOKEN |
Yes | Cloudflare Tunnel token (dashboard access) |
EMAIL_FROM |
Yes | From-address for OTP emails |
EMAIL_TRANSPORT |
No | smtp (default) or saasmail |
SMTP_HOST / SMTP_PORT |
If smtp |
SMTP server (defaults target the bundled Inbucket) |
SAASMAIL_API_KEY / SAASMAIL_API_URL |
If saasmail, or for alerts |
HTTP email API credentials/endpoint |
ALERT_FROM_EMAIL |
No | Sender for alert emails |
RETENTION_INTERVAL_HOURS |
No | How often retention trims expired logs (default 24) |
CF_ACCOUNT_ID / CF_QUEUE_ID / CF_API_TOKEN |
Yes | Queue consumer → Cloudflare Queue access |
LOG_CANNON_INGEST_URL / LOG_CANNON_ADMIN_KEY |
Yes | Ingest Worker URL and admin-scoped key so the dashboard can manage the D1 key registry |
LOG_CANNON_QUEUE_CONSUMER_API_KEY / LOG_CANNON_ALERT_WORKER_API_KEY / LOG_CANNON_RETENTION_WORKER_API_KEY |
No | Per-service ingest keys so the Go workers ship CLEF (see Go service logs); key names should be log-cannon-queue-consumer, log-cannon-alert-worker, log-cannon-retention-worker |
SUPABASE_PULL_PROJECTS / SUPABASE_ACCESS_TOKEN |
No | Supabase projects to pull platform logs from, and a PAT with analytics_logs_read (see Supabase platform logs) |
R2_* |
No | Offsite backup credentials (see Backup & Restore) |
COMPOSE_PROFILES |
No | dev locally to start the Inbucket mailbox |
log-cannon/
├── workers/ # Cloudflare Workers (TypeScript) — edge ingestion
│ └── packages/ingest/ # Unified ingest worker (CLEF, webhook, OTel)
├── go/ship/ # Shared Go CLEF shipper (replace-path dep of the Go services)
├── queue-consumer/ # Go service: pulls CF Queue → parses → ClickHouse
├── dashboard/ # Next.js web UI, REST API, MCP server (reads ClickHouse)
├── alert-worker/ # Go service: threshold alerting
├── retention-worker/ # Go service: per-source retention trimming
├── supabase-pull/ # Go service: Supabase platform logs → ClickHouse (optional)
├── clickhouse/ # Database image + numbered schema init (clickhouse/init)
├── backup/ # Backup/restore scripts with R2 offsite sync
└── docker-compose.yml
- Edge ingestion: Cloudflare Workers (TypeScript) + Queues + D1
- Queue consumer / alert / retention workers: Go 1.22+
- Dashboard / API / MCP: Next.js, React 18, Tailwind CSS
- Storage: ClickHouse
- Infrastructure: Docker Compose, Cloudflare Tunnel, Cloudflare R2 (backups)
Issues and pull requests are welcome. The repo is a monorepo of independent services — see AGENTS.md for a developer-oriented map of how the pieces fit together, how to build and run each one, and the conventions to follow.
MIT — see LICENSE.md.