Skip to content

Scorer declaring only task_span passes construction, then TypeErrors #8676

Description

@feiiiiii5

A scorer that declares only task_span is accepted at construction and then fails on every call.

What happens

validate_scorer_function accepts a scorer with "at least one task_span parameter" — the alternative to having both dataset_item and task_outputs. So this is a valid scorer:

def spans_are_present(task_span=None):
    return score_result.ScoreResult(name="spans_are_present", value=1.0 if task_span else 0.0)

It never runs. Every score() call ends in:

TypeError: spans_are_present() got an unexpected keyword argument 'dataset_item'
Reproduction, root cause, suggested direction

Reproduction

from opik.evaluation.scorers import scorer_wrapper_metric

def spans_are_present(task_span=None):
    return score_result.ScoreResult(name="spans_are_present", value=1.0)

metric = scorer_wrapper_metric.wrap_scorer_functions([spans_are_present], project_name=None)[0]
metric.score(dataset_item={"input": "hi"}, task_outputs={"output": "hello"})

Run against main (2438f12fe2), with only the SDK installed:

  • wrap_scorer_functions succeeds — the ValueError from validate_scorer_function is not raised.
  • metric.score(...) raises TypeError: spans_are_present() got an unexpected keyword argument 'dataset_item'.

Why

ScorerWrapperMetric.score has two ways to reach the scorer:

if needs_task_span and task_span is not None:
    return self.scorer(dataset_item=..., task_outputs=..., task_span=task_span)
return self.scorer(dataset_item=dataset_item, task_outputs=task_outputs)

The second branch is reached whenever no span is bound, and it always passes dataset_item= and task_outputs=. A scorer that takes neither therefore gets two keywords it does not declare.

The two helpers already draw the distinction correctly — has_task_span_in_parameters is True and requires_task_span_argument is False here, because task_span has a default — but that only routes around the "missing required argument" report; nothing filters the keywords before the call.

Impact

A user whose scorer needs nothing but the span cannot re-score or evaluate at all: every item fails with a TypeError from user code rather than the ScoreMethodMissingArguments the tolerance rules know about, so with a strict error_tolerance the run aborts and with a lenient one every score is marked failed for the wrong reason. The docstring for evaluate_experiment says a scorer "whose task_span parameter has a default still runs", so this contradicts documented behaviour.

Suggested direction

I would filter the call by the scorer's own signature — pass only the arguments it actually declares, so a task-span-only scorer is called as scorer(task_span=None). That keeps the existing behaviour for scorers that do declare dataset_item/task_outputs and fixes the ones that do not. An alternative is to reject task-span-only scorers at construction time, which is simpler but breaks callers who have one today and would then get a ValueError instead of a score.

Happy to open a PR either way — please say which you prefer.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions