Skip to content

[FR] Add a tracking server readiness endpoint that checks backend database availability #26333

Description

@youssefcamao

Warning

Before submitting a PR, please make sure that:

  • A maintainer has triaged this issue and applied the ready label
  • This issue has no assignee
  • No duplicate PR exists

PRs not meeting these requirements may be automatically closed.

Willingness to contribute

Yes. I can contribute this feature independently.

Proposal Summary

Add a dedicated readiness endpoint to the MLflow tracking server, for example
GET /ready, that verifies backend database availability through the server's
existing connection pool. Return HTTP 200 when the check succeeds and HTTP 503
when the database is unavailable or the check times out.

Keep the existing /health endpoint's lightweight liveness behavior. Deployers
could use /health for liveness and the new endpoint for database-aware readiness.

Motivation

Add a dedicated readiness endpoint to the MLflow tracking server, for example
GET /ready, that verifies backend database availability through the server's
existing connection pool. Return HTTP 200 when the check succeeds and HTTP 503
when the database is unavailable or the check times out.

Keep the existing /health endpoint's lightweight liveness behavior. Deployers
could use /health for liveness and the new endpoint for database-aware readiness.

Motivation

What is the use case for this feature?

We deploy MLflow tracking servers on Kubernetes with external PostgreSQL
databases. If a running server loses database connectivity, database-backed API
requests fail, while the /health handler continues to return HTTP 200. Using
that endpoint for readiness therefore does not detect this failure.

A separate readiness endpoint would let Kubernetes stop routing traffic to
affected pods while keeping their processes running, then restore traffic when
database connectivity recovers.

Why is this use case valuable to support for MLflow users in general?

Deployers using SQL backend stores could distinguish application responsiveness
from database availability without writing their own database probes. A check
inside MLflow could exercise the connection pool used by actual requests and
reuse established connections.

Why is this use case valuable to support for your project(s) or organization?

We operate multiple MLflow instances with multiple replicas. Our current
readiness workaround launches Python in each container, calls the local
/health endpoint, and opens and closes a separate PostgreSQL connection every
20 seconds.

For example, 50 instances with three replicas would perform 450 fresh database
connections per minute across the fleet. These connections are short-lived, but
the workaround introduces process startup and database connection setup overhead
and requires us to maintain a probe script alongside MLflow.

Why is it currently difficult to achieve this use case?

The /health handler only verifies HTTP responsiveness. Backend store
initialization checks the database during startup, but does not provide a
dedicated readiness signal for database failures after startup.

An external probe has its own database connection lifecycle and can succeed even
when MLflow's own connection pool has a problem. Calling a database-backed
tracking API is another workaround, but ties readiness to that API's
authentication, authorization, and request semantics.

Details

The proposed initial scope is a small, read-only database availability check:

  • Add a separate readiness route, with its name and configuration subject to
    maintainer agreement.
  • For SQLAlchemy backend stores, borrow a connection from the existing engine
    pool, execute a lightweight query such as SELECT 1, and return the connection
    to the pool.
  • Return HTTP 200 on success and HTTP 503 on failure or timeout. Keep the response
    generic; do not expose credentials, connection strings, or database exception
    details.
  • Bound pool acquisition, connection establishment, and query execution time so
    failed checks do not accumulate blocked server workers. A client-side HTTP
    probe timeout alone would not guarantee that the server stops the check.
  • Support the configured static prefix and define how infrastructure probes can
    access the endpoint when authentication is enabled.
  • Test successful checks, database outages, timeouts, and recovery without a
    server restart.

This check would establish basic database responsiveness, not guarantee that
every MLflow operation succeeds or that the database is writable. Artifact
storage checks and client SDK additions are outside the proposed initial scope.

What machine learning domain(s) is this feature request about?

  • domain/genai: LLMs, Agents, and other GenAI-related use cases
  • domain/classical-ml: Traditional machine learning, such as linear regression.
  • domain/deep-learning: Deep learning and neural networks.
  • domain/platform: MLflow platform foundation, not specific to a particular machine learning domain.

What area(s) of MLflow is this feature request about?

  • area/tracking: Tracking Service, tracking client APIs, autologging
  • area/model-registry: Model Registry service, APIs, and the fluent client calls for Model Registry
  • area/scoring: MLflow model serving, deployment tools, Spark UDFs
  • area/evaluation: MLflow model evaluation features, evaluation metrics, and evaluation workflows
  • area/prompt: MLflow prompt engineering features, prompt templates, and prompt management
  • area/tracing: MLflow Tracing features, tracing APIs, and LLM tracing functionality
  • area/gateway: MLflow AI Gateway client APIs, server, and third-party integrations
  • area/projects: MLproject format, project running backends
  • area/uiux: Front-end, user experience, plotting
  • area/docs: MLflow documentation pages

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    AcknowledgedThis issue has been read and acknowledged by the MLflow admins.area/docsDocumentation issuesarea/trackingTracking service, tracking client APIs, autologgingdomain/platformFeature requests about platform foundation, not specific to a problem domainenhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions