Skip to content

[FR] Paired statistical comparison of GenAI evaluation runs (mlflow.genai.compare_evaluations) #26193

Description

@AbdulAliMamnun

Warning

Before submitting a PR, please make sure that:

  • A maintainer has triaged this issue and applied the ready label
  • This issue has no assignee
  • No duplicate PR exists

PRs not meeting these requirements may be automatically closed.

Willingness to contribute

Yes. I can contribute this feature independently.

Proposal Summary

Add mlflow.genai.compare_evaluations(candidate_run_id, baseline_run_id): a paired, per-scorer statistical comparison of two evaluation runs. For each scorer present in both runs it reports the difference, a confidence interval, a significance test selected by the scorer's value type (McNemar for pass/fail feedback, Wilcoxon signed-rank / paired t for numeric scores), an effect size, and the paired sample size. Results are logged back to the candidate run as compare/<scorer>/* metrics plus a per-row deltas artifact, and an assert_improved() helper provides a CI gate meaning "better, and not noise."

Motivation

What is the use case for this feature?

A team evaluates agent v1 and v2 on the same evaluation dataset with mlflow.genai.evaluate() and needs to decide whether v2 is actually better before promoting it.

Why is this use case valuable to support for MLflow users in general?

mlflow.genai.evaluate() reduces each scorer to a point estimate (mean by default; min/max/variance/p90 on request). Comparing two runs is therefore comparing two means with no notion of sample size, pairing, or noise. GenAI evaluation is unusually noisy: scorers are often LLM judges whose outputs vary call to call, and eval sets are small (50–300 rows) because each row costs a model call and a judge call. Under these conditions a 3-point gain in mean correctness is routinely within noise, so teams either ship regressions or block real improvements. MLflow's own evaluation-metrics guide (July 2026) recommends paired significance tests such as Wilcoxon signed-rank before declaring a winner, but the product does not implement that advice. Competing tooling has moved: Azure AI Foundry's evaluation GitHub Action reports confidence intervals per variant and significance tests between variants.

Why is this use case valuable to support for your project(s) or organization?

In a regulated setting (insurance), promoting an LLM application needs a defensible, auditable statement that quality did not regress. "Mean went from 0.71 to 0.76" is not defensible; "paired difference +0.05, 95% CI [-0.01, +0.11], p=0.09 on n=200" is.

Why is it currently difficult to achieve this use case?

The only comparison machinery in the codebase is mlflow.validate_evaluation_results with threshold, min_absolute_change and min_relative_change: deterministic cutoffs applied to point estimates, which cannot distinguish a real 5-point gain on n=500 from a chance 5-point gain on n=30. Doing it in user space means exporting result_df from both runs, re-deriving row identity to pair them, choosing tests by hand, and there is no way to persist the result on the run or gate CI on it. MLflow is well placed to fix this because pairing is free: both runs execute against the same dataset and each evaluation trace carries the record identity (mlflow.eval.requestId), so per-row deltas can be computed exactly, which gives far more statistical power than unpaired tests at these sample sizes.

Details

Proposed shape (happy to write this up as an RFC if it's judged substantial enough):

  • Pairing: join the two runs' evaluation rows on the eval request id; fall back to a stable hash of inputs when ids differ. Unpaired rows are reported, not silently included; fewer than two pairs → insufficient_pairs for that scorer.
  • Test selection (method="auto"): binary values → McNemar (exact for small n) with a Wilson/bootstrap interval on the paired proportion difference; numeric/ordinal → Wilcoxon signed-rank as default with paired t reported alongside, percentile bootstrap CI on the mean paired delta, Cohen's d_z effect size. Ties counted and reported.
  • Direction: greater_is_better from the scorer definition, overridable per scorer.
  • Persistence: metrics compare/<scorer>/{diff,ci_low,ci_high,p_value,effect_size,n_paired} and tag mlflow.compare.baselineRunId on the candidate run; one JSON artifact of per-row deltas. No new tables, no server changes.
  • Gating: ComparisonResult.assert_improved(scorers, alpha) raises unless every listed scorer moved in the better direction with p < alpha; composes with existing threshold validation.
  • Dependencies: scipy only (already core).
  • Out of scope for v1: multiple-comparison correction across scorers (documented caveat, natural follow-on), comparing more than two runs, a dedicated UI view, modelling judge non-determinism (docs will recommend an A/A run as calibration).

Rough footprint: mlflow/genai/evaluation/comparison.py, ~400–600 lines including tests, plus a docs page "Comparing evaluation runs".

What machine learning domain(s) is this feature request about?

  • domain/genai: LLMs, Agents, and other GenAI-related use cases
  • domain/classical-ml: Traditional machine learning, such as linear regression.
  • domain/deep-learning: Deep learning and neural networks.
  • domain/platform: MLflow platform foundation, not specific to a particular machine learning domain.

What area(s) of MLflow is this feature request about?

  • area/tracking: Tracking Service, tracking client APIs, autologging
  • area/model-registry: Model Registry service, APIs, and the fluent client calls for Model Registry
  • area/scoring: MLflow model serving, deployment tools, Spark UDFs
  • area/evaluation: MLflow model evaluation features, evaluation metrics, and evaluation workflows
  • area/prompt: MLflow prompt engineering features, prompt templates, and prompt management
  • area/tracing: MLflow Tracing features, tracing APIs, and LLM tracing functionality
  • area/gateway: MLflow AI Gateway client APIs, server, and third-party integrations
  • area/projects: MLproject format, project running backends
  • area/uiux: Front-end, user experience, plotting
  • area/docs: MLflow documentation pages

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

area/evaluationMLflow Evaluationdomain/genaiFeature requests related to GenAI/LLM use casesenhancementNew feature or requesthas-closing-prThis issue has a closing PRreadyTriaged and ready for implementation

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions