You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
A maintainer has triaged this issue and applied the ready label
This issue has no assignee
No duplicate PR exists
PRs not meeting these requirements may be automatically closed.
Willingness to contribute
Yes. I can contribute this feature independently.
Proposal Summary
Add a dedicated readiness endpoint to the MLflow tracking server, for example GET /ready, that verifies backend database availability through the server's
existing connection pool. Return HTTP 200 when the check succeeds and HTTP 503
when the database is unavailable or the check times out.
Keep the existing /health endpoint's lightweight liveness behavior. Deployers
could use /health for liveness and the new endpoint for database-aware readiness.
Motivation
Add a dedicated readiness endpoint to the MLflow tracking server, for example GET /ready, that verifies backend database availability through the server's
existing connection pool. Return HTTP 200 when the check succeeds and HTTP 503
when the database is unavailable or the check times out.
Keep the existing /health endpoint's lightweight liveness behavior. Deployers
could use /health for liveness and the new endpoint for database-aware readiness.
Motivation
What is the use case for this feature?
We deploy MLflow tracking servers on Kubernetes with external PostgreSQL
databases. If a running server loses database connectivity, database-backed API
requests fail, while the /health handler continues to return HTTP 200. Using
that endpoint for readiness therefore does not detect this failure.
A separate readiness endpoint would let Kubernetes stop routing traffic to
affected pods while keeping their processes running, then restore traffic when
database connectivity recovers.
Why is this use case valuable to support for MLflow users in general?
Deployers using SQL backend stores could distinguish application responsiveness
from database availability without writing their own database probes. A check
inside MLflow could exercise the connection pool used by actual requests and
reuse established connections.
Why is this use case valuable to support for your project(s) or organization?
We operate multiple MLflow instances with multiple replicas. Our current
readiness workaround launches Python in each container, calls the local /health endpoint, and opens and closes a separate PostgreSQL connection every
20 seconds.
For example, 50 instances with three replicas would perform 450 fresh database
connections per minute across the fleet. These connections are short-lived, but
the workaround introduces process startup and database connection setup overhead
and requires us to maintain a probe script alongside MLflow.
Why is it currently difficult to achieve this use case?
The /health handler only verifies HTTP responsiveness. Backend store
initialization checks the database during startup, but does not provide a
dedicated readiness signal for database failures after startup.
An external probe has its own database connection lifecycle and can succeed even
when MLflow's own connection pool has a problem. Calling a database-backed
tracking API is another workaround, but ties readiness to that API's
authentication, authorization, and request semantics.
Details
The proposed initial scope is a small, read-only database availability check:
Add a separate readiness route, with its name and configuration subject to
maintainer agreement.
For SQLAlchemy backend stores, borrow a connection from the existing engine
pool, execute a lightweight query such as SELECT 1, and return the connection
to the pool.
Return HTTP 200 on success and HTTP 503 on failure or timeout. Keep the response
generic; do not expose credentials, connection strings, or database exception
details.
Bound pool acquisition, connection establishment, and query execution time so
failed checks do not accumulate blocked server workers. A client-side HTTP
probe timeout alone would not guarantee that the server stops the check.
Support the configured static prefix and define how infrastructure probes can
access the endpoint when authentication is enabled.
Test successful checks, database outages, timeouts, and recovery without a
server restart.
This check would establish basic database responsiveness, not guarantee that
every MLflow operation succeeds or that the database is writable. Artifact
storage checks and client SDK additions are outside the proposed initial scope.
What machine learning domain(s) is this feature request about?
domain/genai: LLMs, Agents, and other GenAI-related use cases
domain/classical-ml: Traditional machine learning, such as linear regression.
domain/deep-learning: Deep learning and neural networks.
domain/platform: MLflow platform foundation, not specific to a particular machine learning domain.
What area(s) of MLflow is this feature request about?
Warning
Before submitting a PR, please make sure that:
readylabelPRs not meeting these requirements may be automatically closed.
Willingness to contribute
Yes. I can contribute this feature independently.
Proposal Summary
Add a dedicated readiness endpoint to the MLflow tracking server, for example
GET /ready, that verifies backend database availability through the server'sexisting connection pool. Return HTTP 200 when the check succeeds and HTTP 503
when the database is unavailable or the check times out.
Keep the existing
/healthendpoint's lightweight liveness behavior. Deployerscould use
/healthfor liveness and the new endpoint for database-aware readiness.Motivation
Add a dedicated readiness endpoint to the MLflow tracking server, for example
GET /ready, that verifies backend database availability through the server'sexisting connection pool. Return HTTP 200 when the check succeeds and HTTP 503
when the database is unavailable or the check times out.
Keep the existing
/healthendpoint's lightweight liveness behavior. Deployerscould use
/healthfor liveness and the new endpoint for database-aware readiness.Motivation
What is the use case for this feature?
We deploy MLflow tracking servers on Kubernetes with external PostgreSQL
databases. If a running server loses database connectivity, database-backed API
requests fail, while the
/healthhandler continues to return HTTP 200. Usingthat endpoint for readiness therefore does not detect this failure.
A separate readiness endpoint would let Kubernetes stop routing traffic to
affected pods while keeping their processes running, then restore traffic when
database connectivity recovers.
Why is this use case valuable to support for MLflow users in general?
Deployers using SQL backend stores could distinguish application responsiveness
from database availability without writing their own database probes. A check
inside MLflow could exercise the connection pool used by actual requests and
reuse established connections.
Why is this use case valuable to support for your project(s) or organization?
We operate multiple MLflow instances with multiple replicas. Our current
readiness workaround launches Python in each container, calls the local
/healthendpoint, and opens and closes a separate PostgreSQL connection every20 seconds.
For example, 50 instances with three replicas would perform 450 fresh database
connections per minute across the fleet. These connections are short-lived, but
the workaround introduces process startup and database connection setup overhead
and requires us to maintain a probe script alongside MLflow.
Why is it currently difficult to achieve this use case?
The
/healthhandler only verifies HTTP responsiveness. Backend storeinitialization checks the database during startup, but does not provide a
dedicated readiness signal for database failures after startup.
An external probe has its own database connection lifecycle and can succeed even
when MLflow's own connection pool has a problem. Calling a database-backed
tracking API is another workaround, but ties readiness to that API's
authentication, authorization, and request semantics.
Details
The proposed initial scope is a small, read-only database availability check:
maintainer agreement.
pool, execute a lightweight query such as
SELECT 1, and return the connectionto the pool.
generic; do not expose credentials, connection strings, or database exception
details.
failed checks do not accumulate blocked server workers. A client-side HTTP
probe timeout alone would not guarantee that the server stops the check.
access the endpoint when authentication is enabled.
server restart.
This check would establish basic database responsiveness, not guarantee that
every MLflow operation succeeds or that the database is writable. Artifact
storage checks and client SDK additions are outside the proposed initial scope.
What machine learning domain(s) is this feature request about?
domain/genai: LLMs, Agents, and other GenAI-related use casesdomain/classical-ml: Traditional machine learning, such as linear regression.domain/deep-learning: Deep learning and neural networks.domain/platform: MLflow platform foundation, not specific to a particular machine learning domain.What area(s) of MLflow is this feature request about?
area/tracking: Tracking Service, tracking client APIs, autologgingarea/model-registry: Model Registry service, APIs, and the fluent client calls for Model Registryarea/scoring: MLflow model serving, deployment tools, Spark UDFsarea/evaluation: MLflow model evaluation features, evaluation metrics, and evaluation workflowsarea/prompt: MLflow prompt engineering features, prompt templates, and prompt managementarea/tracing: MLflow Tracing features, tracing APIs, and LLM tracing functionalityarea/gateway: MLflow AI Gateway client APIs, server, and third-party integrationsarea/projects: MLproject format, project running backendsarea/uiux: Front-end, user experience, plottingarea/docs: MLflow documentation pages