You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
A maintainer has triaged this issue and applied the ready label
This issue has no assignee
No duplicate PR exists
PRs not meeting these requirements may be automatically closed.
Willingness to contribute
Yes. I can contribute this feature independently.
Proposal Summary
Add mlflow.genai.compare_evaluations(candidate_run_id, baseline_run_id): a paired, per-scorer statistical comparison of two evaluation runs. For each scorer present in both runs it reports the difference, a confidence interval, a significance test selected by the scorer's value type (McNemar for pass/fail feedback, Wilcoxon signed-rank / paired t for numeric scores), an effect size, and the paired sample size. Results are logged back to the candidate run as compare/<scorer>/* metrics plus a per-row deltas artifact, and an assert_improved() helper provides a CI gate meaning "better, and not noise."
Motivation
What is the use case for this feature?
A team evaluates agent v1 and v2 on the same evaluation dataset with mlflow.genai.evaluate() and needs to decide whether v2 is actually better before promoting it.
Why is this use case valuable to support for MLflow users in general?
mlflow.genai.evaluate() reduces each scorer to a point estimate (mean by default; min/max/variance/p90 on request). Comparing two runs is therefore comparing two means with no notion of sample size, pairing, or noise. GenAI evaluation is unusually noisy: scorers are often LLM judges whose outputs vary call to call, and eval sets are small (50–300 rows) because each row costs a model call and a judge call. Under these conditions a 3-point gain in mean correctness is routinely within noise, so teams either ship regressions or block real improvements. MLflow's own evaluation-metrics guide (July 2026) recommends paired significance tests such as Wilcoxon signed-rank before declaring a winner, but the product does not implement that advice. Competing tooling has moved: Azure AI Foundry's evaluation GitHub Action reports confidence intervals per variant and significance tests between variants.
Why is this use case valuable to support for your project(s) or organization?
In a regulated setting (insurance), promoting an LLM application needs a defensible, auditable statement that quality did not regress. "Mean went from 0.71 to 0.76" is not defensible; "paired difference +0.05, 95% CI [-0.01, +0.11], p=0.09 on n=200" is.
Why is it currently difficult to achieve this use case?
The only comparison machinery in the codebase is mlflow.validate_evaluation_results with threshold, min_absolute_change and min_relative_change: deterministic cutoffs applied to point estimates, which cannot distinguish a real 5-point gain on n=500 from a chance 5-point gain on n=30. Doing it in user space means exporting result_df from both runs, re-deriving row identity to pair them, choosing tests by hand, and there is no way to persist the result on the run or gate CI on it. MLflow is well placed to fix this because pairing is free: both runs execute against the same dataset and each evaluation trace carries the record identity (mlflow.eval.requestId), so per-row deltas can be computed exactly, which gives far more statistical power than unpaired tests at these sample sizes.
Details
Proposed shape (happy to write this up as an RFC if it's judged substantial enough):
Pairing: join the two runs' evaluation rows on the eval request id; fall back to a stable hash of inputs when ids differ. Unpaired rows are reported, not silently included; fewer than two pairs → insufficient_pairs for that scorer.
Test selection (method="auto"): binary values → McNemar (exact for small n) with a Wilson/bootstrap interval on the paired proportion difference; numeric/ordinal → Wilcoxon signed-rank as default with paired t reported alongside, percentile bootstrap CI on the mean paired delta, Cohen's d_z effect size. Ties counted and reported.
Direction:greater_is_better from the scorer definition, overridable per scorer.
Persistence: metrics compare/<scorer>/{diff,ci_low,ci_high,p_value,effect_size,n_paired} and tag mlflow.compare.baselineRunId on the candidate run; one JSON artifact of per-row deltas. No new tables, no server changes.
Gating:ComparisonResult.assert_improved(scorers, alpha) raises unless every listed scorer moved in the better direction with p < alpha; composes with existing threshold validation.
Dependencies: scipy only (already core).
Out of scope for v1: multiple-comparison correction across scorers (documented caveat, natural follow-on), comparing more than two runs, a dedicated UI view, modelling judge non-determinism (docs will recommend an A/A run as calibration).
Rough footprint: mlflow/genai/evaluation/comparison.py, ~400–600 lines including tests, plus a docs page "Comparing evaluation runs".
What machine learning domain(s) is this feature request about?
domain/genai: LLMs, Agents, and other GenAI-related use cases
domain/classical-ml: Traditional machine learning, such as linear regression.
domain/deep-learning: Deep learning and neural networks.
domain/platform: MLflow platform foundation, not specific to a particular machine learning domain.
What area(s) of MLflow is this feature request about?
Warning
Before submitting a PR, please make sure that:
readylabelPRs not meeting these requirements may be automatically closed.
Willingness to contribute
Yes. I can contribute this feature independently.
Proposal Summary
Add
mlflow.genai.compare_evaluations(candidate_run_id, baseline_run_id): a paired, per-scorer statistical comparison of two evaluation runs. For each scorer present in both runs it reports the difference, a confidence interval, a significance test selected by the scorer's value type (McNemar for pass/fail feedback, Wilcoxon signed-rank / paired t for numeric scores), an effect size, and the paired sample size. Results are logged back to the candidate run ascompare/<scorer>/*metrics plus a per-row deltas artifact, and anassert_improved()helper provides a CI gate meaning "better, and not noise."Motivation
A team evaluates agent v1 and v2 on the same evaluation dataset with
mlflow.genai.evaluate()and needs to decide whether v2 is actually better before promoting it.mlflow.genai.evaluate()reduces each scorer to a point estimate (mean by default; min/max/variance/p90 on request). Comparing two runs is therefore comparing two means with no notion of sample size, pairing, or noise. GenAI evaluation is unusually noisy: scorers are often LLM judges whose outputs vary call to call, and eval sets are small (50–300 rows) because each row costs a model call and a judge call. Under these conditions a 3-point gain in mean correctness is routinely within noise, so teams either ship regressions or block real improvements. MLflow's own evaluation-metrics guide (July 2026) recommends paired significance tests such as Wilcoxon signed-rank before declaring a winner, but the product does not implement that advice. Competing tooling has moved: Azure AI Foundry's evaluation GitHub Action reports confidence intervals per variant and significance tests between variants.In a regulated setting (insurance), promoting an LLM application needs a defensible, auditable statement that quality did not regress. "Mean went from 0.71 to 0.76" is not defensible; "paired difference +0.05, 95% CI [-0.01, +0.11], p=0.09 on n=200" is.
The only comparison machinery in the codebase is
mlflow.validate_evaluation_resultswiththreshold,min_absolute_changeandmin_relative_change: deterministic cutoffs applied to point estimates, which cannot distinguish a real 5-point gain on n=500 from a chance 5-point gain on n=30. Doing it in user space means exportingresult_dffrom both runs, re-deriving row identity to pair them, choosing tests by hand, and there is no way to persist the result on the run or gate CI on it. MLflow is well placed to fix this because pairing is free: both runs execute against the same dataset and each evaluation trace carries the record identity (mlflow.eval.requestId), so per-row deltas can be computed exactly, which gives far more statistical power than unpaired tests at these sample sizes.Details
Proposed shape (happy to write this up as an RFC if it's judged substantial enough):
inputswhen ids differ. Unpaired rows are reported, not silently included; fewer than two pairs →insufficient_pairsfor that scorer.method="auto"): binary values → McNemar (exact for small n) with a Wilson/bootstrap interval on the paired proportion difference; numeric/ordinal → Wilcoxon signed-rank as default with paired t reported alongside, percentile bootstrap CI on the mean paired delta, Cohen's d_z effect size. Ties counted and reported.greater_is_betterfrom the scorer definition, overridable per scorer.compare/<scorer>/{diff,ci_low,ci_high,p_value,effect_size,n_paired}and tagmlflow.compare.baselineRunIdon the candidate run; one JSON artifact of per-row deltas. No new tables, no server changes.ComparisonResult.assert_improved(scorers, alpha)raises unless every listed scorer moved in the better direction with p < alpha; composes with existing threshold validation.Rough footprint:
mlflow/genai/evaluation/comparison.py, ~400–600 lines including tests, plus a docs page "Comparing evaluation runs".What machine learning domain(s) is this feature request about?
domain/genai: LLMs, Agents, and other GenAI-related use casesdomain/classical-ml: Traditional machine learning, such as linear regression.domain/deep-learning: Deep learning and neural networks.domain/platform: MLflow platform foundation, not specific to a particular machine learning domain.What area(s) of MLflow is this feature request about?
area/tracking: Tracking Service, tracking client APIs, autologgingarea/model-registry: Model Registry service, APIs, and the fluent client calls for Model Registryarea/scoring: MLflow model serving, deployment tools, Spark UDFsarea/evaluation: MLflow model evaluation features, evaluation metrics, and evaluation workflowsarea/prompt: MLflow prompt engineering features, prompt templates, and prompt managementarea/tracing: MLflow Tracing features, tracing APIs, and LLM tracing functionalityarea/gateway: MLflow AI Gateway client APIs, server, and third-party integrationsarea/projects: MLproject format, project running backendsarea/uiux: Front-end, user experience, plottingarea/docs: MLflow documentation pages