You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
A maintainer has triaged this issue and applied the ready label
This issue has no assignee
No duplicate PR exists
PRs not meeting these requirements may be automatically closed.
Issues Policy acknowledgement
I have read and agree to submit bug reports in accordance with the issues policy
Where did you encounter this bug?
Local machine
MLflow version
Client: source checkout of master, 26a2cfe8cb832b446a6b0fe1e99119cb6e5dcead, 3.16.2.dev0
Tracking server: not required for the minimal reproduction; the end-to-end check uses a local SQLite tracking URI.
System information
OS: macOS, arm64
Python: 3.12.14
NumPy: 2.5.3
pandas: 3.0.6
scikit-learn: 1.9.1
Installed from the repository with its frozen uv lockfile and pytest dependency group. No API keys, GPU or model downloads are required.
Describe the problem
precision_at_k, recall_at_k, and ndcg_at_k reject document-ID arrays with dtype np.int32 (also np.int8 and np.int16) as non-integer input. The same IDs in Python lists or np.int64 arrays produce the expected scores.
The retriever evaluation documentation supports string/integer document IDs in lists or NumPy arrays. Changing the storage width of an integer ID should not remove the evaluation metrics. In an end-to-end mlflow.models.evaluate(..., model_type="retriever") check, int64 yields all nine score aggregations; identical int32 input yields an empty result.metrics dictionary, with skip warnings.
The shared validator _validate_array_like_id_data uses np.issubdtype(value.dtype, int). NumPy treats int as the concrete platform-default integer dtype here, rather than the integer family. On this machine:
The NumPy documentation demonstrates np.integer for recognizing an int32 array. A proposed narrow fix is to use the integer family in the ndarray branch, preserving the existing string branch and rejection of floating-point IDs. No public API or dependency change is needed. I can contribute the fix and regression tests once this issue is triaged and the approach is accepted.
These are legacy metrics deprecated since MLflow 3.4.0; this report concerns their currently documented behavior and does not propose a new evaluation API. The existing NDCG issue #24541 / PR #26125 concerns normalization when fewer than k documents are retrieved. This reproduction retrieves exactly k documents and fails before any score calculation, so it is separate.
Prepared and tested with the help of Codex.
Tracking information
The minimal metric reproduction does not access a tracking backend. The end-to-end check used a temporary local SQLite database and local artifact files, without a tracking server.
No exception is raised; the metrics return None. The validator logs this warning for each metric (shown for precision):
WARNING mlflow.metrics.metric_definitions: Cannot calculate metric 'precision_at_k' for non-arraylike of string or int inputs. Non-arraylike of strings/ints found for the column specified by the `predictions` parameter or the model output column on row 0, value [1 2]. Skipping metric logging.
A separate regression suite: 18 failed / 18 passed. The failures cover int8, int16, and int32 in predictions or targets for each of the three metrics; all fail because the result is None. Controls confirm int64 and string arrays work, and floating-point arrays remain rejected. The suite compares both per-row scores and aggregate results with Python-list inputs.
End-to-end local evaluation: int64 produces nine aggregations; identical int32 produces {}.
Warning
Before submitting a PR, please make sure that:
readylabelPRs not meeting these requirements may be automatically closed.
Issues Policy acknowledgement
Where did you encounter this bug?
Local machine
MLflow version
26a2cfe8cb832b446a6b0fe1e99119cb6e5dcead,3.16.2.dev0System information
Describe the problem
precision_at_k,recall_at_k, andndcg_at_kreject document-ID arrays with dtypenp.int32(alsonp.int8andnp.int16) as non-integer input. The same IDs in Python lists ornp.int64arrays produce the expected scores.The retriever evaluation documentation supports string/integer document IDs in lists or NumPy arrays. Changing the storage width of an integer ID should not remove the evaluation metrics. In an end-to-end
mlflow.models.evaluate(..., model_type="retriever")check,int64yields all nine score aggregations; identicalint32input yields an emptyresult.metricsdictionary, with skip warnings.The shared validator
_validate_array_like_id_datausesnp.issubdtype(value.dtype, int). NumPy treatsintas the concrete platform-default integer dtype here, rather than the integer family. On this machine:The NumPy documentation demonstrates
np.integerfor recognizing anint32array. A proposed narrow fix is to use the integer family in the ndarray branch, preserving the existing string branch and rejection of floating-point IDs. No public API or dependency change is needed. I can contribute the fix and regression tests once this issue is triaged and the approach is accepted.These are legacy metrics deprecated since MLflow 3.4.0; this report concerns their currently documented behavior and does not propose a new evaluation API. The existing NDCG issue #24541 / PR #26125 concerns normalization when fewer than k documents are retrieved. This reproduction retrieves exactly k documents and fails before any score calculation, so it is separate.
Prepared and tested with the help of Codex.
Tracking information
The minimal metric reproduction does not access a tracking backend. The end-to-end check used a temporary local SQLite database and local artifact files, without a tracking server.
Code to reproduce issue
Expected:
[1.0]for both dtypes and all three metrics.Actual:
Stack trace
No exception is raised; the metrics return
None. The validator logs this warning for each metric (shown for precision):Other info / logs
Validation on the unmodified master checkout:
int8,int16, andint32in predictions or targets for each of the three metrics; all fail because the result isNone. Controls confirmint64and string arrays work, and floating-point arrays remain rejected. The suite compares both per-row scores and aggregate results with Python-list inputs.int64produces nine aggregations; identicalint32produces{}.Minimal proposed regression test:
What component(s) does this bug affect?
area/tracking: Tracking Service, tracking client APIs, autologgingarea/model-registry: Model Registry service, APIs, and the fluent client calls for Model Registryarea/scoring: MLflow model serving, deployment tools, Spark UDFsarea/evaluation: MLflow model evaluation features, evaluation metrics, and evaluation workflowsarea/prompt: MLflow prompt engineering features, prompt templates, and prompt managementarea/tracing: MLflow Tracing features, tracing APIs, and LLM tracing functionalityarea/gateway: MLflow AI Gateway client APIs, server, and third-party integrationsarea/projects: MLproject format, project running backendsarea/uiux: Front-end, user experience, plottingarea/docs: MLflow documentation pages