Skip to content

[BUG] Retriever metrics reject NumPy int32 document IDs and omit all evaluation scores #26321

Description

@Marcuswang0824

Warning

Before submitting a PR, please make sure that:

  • A maintainer has triaged this issue and applied the ready label
  • This issue has no assignee
  • No duplicate PR exists

PRs not meeting these requirements may be automatically closed.

Issues Policy acknowledgement

  • I have read and agree to submit bug reports in accordance with the issues policy

Where did you encounter this bug?

Local machine

MLflow version

  • Client: source checkout of master, 26a2cfe8cb832b446a6b0fe1e99119cb6e5dcead, 3.16.2.dev0
  • Tracking server: not required for the minimal reproduction; the end-to-end check uses a local SQLite tracking URI.

System information

  • OS: macOS, arm64
  • Python: 3.12.14
  • NumPy: 2.5.3
  • pandas: 3.0.6
  • scikit-learn: 1.9.1
  • Installed from the repository with its frozen uv lockfile and pytest dependency group. No API keys, GPU or model downloads are required.

Describe the problem

precision_at_k, recall_at_k, and ndcg_at_k reject document-ID arrays with dtype np.int32 (also np.int8 and np.int16) as non-integer input. The same IDs in Python lists or np.int64 arrays produce the expected scores.

The retriever evaluation documentation supports string/integer document IDs in lists or NumPy arrays. Changing the storage width of an integer ID should not remove the evaluation metrics. In an end-to-end mlflow.models.evaluate(..., model_type="retriever") check, int64 yields all nine score aggregations; identical int32 input yields an empty result.metrics dictionary, with skip warnings.

The shared validator _validate_array_like_id_data uses np.issubdtype(value.dtype, int). NumPy treats int as the concrete platform-default integer dtype here, rather than the integer family. On this machine:

dtype   issubdtype(dtype, int)  issubdtype(dtype, np.integer)
int8    False                  True
int16   False                  True
int32   False                  True
int64   True                   True

The NumPy documentation demonstrates np.integer for recognizing an int32 array. A proposed narrow fix is to use the integer family in the ndarray branch, preserving the existing string branch and rejection of floating-point IDs. No public API or dependency change is needed. I can contribute the fix and regression tests once this issue is triaged and the approach is accepted.

These are legacy metrics deprecated since MLflow 3.4.0; this report concerns their currently documented behavior and does not propose a new evaluation API. The existing NDCG issue #24541 / PR #26125 concerns normalization when fewer than k documents are retrieved. This reproduction retrieves exactly k documents and fails before any score calculation, so it is separate.

Prepared and tested with the help of Codex.

Tracking information

The minimal metric reproduction does not access a tracking backend. The end-to-end check used a temporary local SQLite database and local artifact files, without a tracking server.

Code to reproduce issue

import mlflow
import numpy as np
import pandas as pd
from mlflow.metrics import ndcg_at_k, precision_at_k, recall_at_k

print("MLflow:", mlflow.__version__)
for metric_factory in (precision_at_k, recall_at_k, ndcg_at_k):
    metric = metric_factory(2)
    for dtype in (np.int64, np.int32):
        predictions = pd.Series([np.array([1, 2], dtype=dtype)])
        targets = pd.Series([np.array([1, 2], dtype=dtype)])
        result = metric.eval_fn(predictions, targets)
        print(metric_factory.__name__, np.dtype(dtype),
              None if result is None else result.scores)

Expected: [1.0] for both dtypes and all three metrics.

Actual:

precision_at_k int64 [1.0]
precision_at_k int32 None
recall_at_k int64 [1.0]
recall_at_k int32 None
ndcg_at_k int64 [1.0]
ndcg_at_k int32 None

Stack trace

No exception is raised; the metrics return None. The validator logs this warning for each metric (shown for precision):

WARNING mlflow.metrics.metric_definitions: Cannot calculate metric 'precision_at_k' for non-arraylike of string or int inputs. Non-arraylike of strings/ints found for the column specified by the `predictions` parameter or the model output column on row 0, value [1 2]. Skipping metric logging.

Other info / logs

Validation on the unmodified master checkout:

  • Existing retrieval metric tests: 11 passed / 29 deselected. Existing deprecation warnings are emitted.
  • A separate regression suite: 18 failed / 18 passed. The failures cover int8, int16, and int32 in predictions or targets for each of the three metrics; all fail because the result is None. Controls confirm int64 and string arrays work, and floating-point arrays remain rejected. The suite compares both per-row scores and aggregate results with Python-list inputs.
  • End-to-end local evaluation: int64 produces nine aggregations; identical int32 produces {}.

Minimal proposed regression test:

import mlflow
import numpy as np
import pandas as pd
import pytest
from mlflow.metrics import ndcg_at_k, precision_at_k, recall_at_k

@pytest.mark.parametrize("factory", [precision_at_k, recall_at_k, ndcg_at_k])
def test_retriever_accepts_int32_document_ids(factory):
    predictions = pd.Series([np.array([1, 2], dtype=np.int32)])
    targets = pd.Series([np.array([1, 2], dtype=np.int32)])
    result = factory(2).eval_fn(predictions, targets)
    assert result is not None
    assert result.scores == pytest.approx([1.0])

What component(s) does this bug affect?

  • area/tracking: Tracking Service, tracking client APIs, autologging
  • area/model-registry: Model Registry service, APIs, and the fluent client calls for Model Registry
  • area/scoring: MLflow model serving, deployment tools, Spark UDFs
  • area/evaluation: MLflow model evaluation features, evaluation metrics, and evaluation workflows
  • area/prompt: MLflow prompt engineering features, prompt templates, and prompt management
  • area/tracing: MLflow Tracing features, tracing APIs, and LLM tracing functionality
  • area/gateway: MLflow AI Gateway client APIs, server, and third-party integrations
  • area/projects: MLproject format, project running backends
  • area/uiux: Front-end, user experience, plotting
  • area/docs: MLflow documentation pages

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

AcknowledgedThis issue has been read and acknowledged by the MLflow admins.bugSomething isn't workingreadyTriaged and ready for implementation

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions