Skip to content

compute_pareto_efficiency crashes with "F" returns in multi-objective CBO #366

Description

@pbalapra

Summary

When a run function returns "F" (the documented failure convention) in multi-objective CBO, compute_pareto_efficiency() crashes with ValueError: array must not contain infs or NaNs. The failure filter only checks the first objective column, so "F" values in other columns slip through as NaN.

Reproducer

from deephyper.evaluator import Evaluator, RunningJob
from deephyper.hpo import CBO, HpProblem

problem = HpProblem()
problem.add_hyperparameter((0.0, 1.0), "x")

def run(job: RunningJob):
    x = job.parameters["x"]
    if x < 0.5:
        return "F"  # fail half the configs
    return {"objective": (x, -x)}

evaluator = Evaluator.create(run, method="thread")
search = CBO(problem, surrogate_model="ET", acq_func="UCB",
             moo_scalarization_strategy="Chebyshev")
results = search.search(evaluator, max_evals=10)  # crashes

Traceback

File "deephyper/hpo/_search.py", line 234, in compute_pareto_efficiency
    mask_pareto_front = non_dominated_set(objectives)
File "deephyper/skopt/moo/_pf.py", line 103, in non_dominated_set
    y = np.asarray_chkfinite(y)
ValueError: array must not contain infs or NaNs

Root Cause

In _search.py:232, compute_pareto_efficiency calls:

_, mask_no_failures = get_mask_of_rows_without_failures(df, objective_columns[0])

This only checks objective_columns[0] for "F" strings. In multi-objective mode, when "F" is returned, DeepHyper stores the failure marker in each objective column. However, pandas may infer mixed dtypes per column — if objective_0 happens to be numeric (e.g., when some evals succeed), get_mask_of_rows_without_failures takes the else branch (isinstance(x, float)) and correctly filters failures. But if objective_1 or objective_2 columns contain "F" strings that weren't checked, the subsequent .astype(float) on line 233 silently converts them to NaN:

objectives = -df.loc[mask_no_failures, objective_columns].values.astype(float)

Then non_dominated_set() calls np.asarray_chkfinite(y), which rejects NaN.

The issue is that get_mask_of_rows_without_failures is called on only the first objective column, not all of them.

Suggested Fix

Check all objective columns and combine the masks:

# _search.py, compute_pareto_efficiency(), around line 232
# Before (buggy):
_, mask_no_failures = get_mask_of_rows_without_failures(df, objective_columns[0])

# After (fixed):
mask_no_failures = np.ones(len(df), dtype=bool)
for col in objective_columns:
    _, col_mask = get_mask_of_rows_without_failures(df, col)
    mask_no_failures &= col_mask

Environment

  • DeepHyper 0.13.2 (latest on PyPI)
  • Python 3.11
  • Multi-objective CBO with Chebyshev scalarization

Workaround

Return a numeric penalty tuple instead of "F":

return {"objective": (-1.0, -1.0, -1.0)}  # instead of return "F"

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions