Skip to content

feat: Gotenberg 9 - #1487

Open
gulien wants to merge 12 commits into
mainfrom
next
Open

gulien wants to merge 12 commits into
mainfrom
next

Conversation

@gulien

@gulien gulien commented Mar 7, 2026 •

Copy link
Copy Markdown
Collaborator

This PR is currently outdated.

@gulien gulien added the enhancement New feature or request label Mar 7, 2026
@gulien gulien changed the title Gotenberg 9 feat: Gotenberg 9 Mar 8, 2026
@jbdelhommeau

Copy link
Copy Markdown

Feature Request: Richer Prometheus Metrics for Better Observability

Hi @gulien 👋

First, thank you for the work on the Prometheus module — having
chromium_requests_queue_size and chromium_restarts_count is a solid
foundation for Kubernetes-based deployments.

That said, we've been running Gotenberg in production at scale (multiple
clusters, HPA-managed, scraped via Datadog OpenMetrics) and we keep hitting
the same wall: the current metrics only reflect internal state, not
conversion behavior
. This makes it impossible to define SLOs, build
meaningful alerts, or make data-driven scaling decisions.

What's missing in practice

Dimension Current coverage
Latency / SLOs ❌ None
Throughput ❌ None
Error rate ❌ None
Active requests ❌ None
PDF output size ❌ None
HTTP API layer ❌ None

Proposed metrics

All metrics below would live in the existing chromium (and libreoffice)
modules, following the same MetricsProvider interface pattern already in
place.


P0 — Critical for SLOs and alerting

chromium_conversion_duration_seconds · Histogram

Duration of each HTML-to-PDF conversion, labeled by status
(success | error | timeout).

Suggested buckets: 0.5, 1, 2, 5, 10, 30, 60

Without this, there is no way to know whether Gotenberg is slow or fast,
degrading or stable. This single metric unlocks p95/p99 SLOs, latency
alerts, and meaningful dashboards.

chromium_requests_total · Counter · label: status

Total number of conversion requests processed.

Enables error rate calculation:
rate(chromium_requests_total{status="error"}[5m]) / rate(chromium_requests_total[5m])


P1 — Operational clarity

chromium_active_requests · Gauge

Number of conversions currently in progress (as opposed to queued).

With maxQueueSize capped, it's critical to distinguish "waiting in
queue" from "actively being processed by Chromium". Today both states
are invisible.

chromium_queue_wait_duration_seconds · Histogram · label: status

Time a request spends waiting in the queue before processing starts.

Allows root-cause analysis: is the SLO broken because Chromium is slow,
or because requests wait too long before even starting?

chromium_errors_total · Counter · label: reason

(timeout | page_crash | context_cancelled | invalid_input
| chromium_unavailable)

Distinguishes client errors from infrastructure errors. Essential for
targeted alerting and on-call triage.


P2 — Capacity planning

chromium_pdf_output_size_bytes · Histogram

Size distribution of generated PDF files.

Useful for correlating large PDFs with latency spikes, and for
bandwidth/storage capacity planning.


Why this matters beyond dashboards

With chromium_active_requests + chromium_conversion_duration_seconds,
operators can replace CPU/memory-based HPA with queue-depth and latency-
driven autoscaling
(e.g. via KEDA), which is far more relevant for a
conversion workload where resource usage varies wildly per document.

# Example: scale on queue depth instead of CPU
triggers:
  - type: prometheus
    metadata:
      query: gotenberg_chromium_requests_queue_size > 2
      threshold: "1"

Implementation notes

All proposed metrics are straightforward float64 reads — no lock
contention, no sampling complexity. They fit cleanly into the existing
Metrics() method pattern used by the Chromium module today.

We'd be happy to contribute a PR if this direction sounds good to you.

Thanks for considering it!

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants