A scorer that declares only task_span is accepted at construction and then fails on every call.
What happens
validate_scorer_function accepts a scorer with "at least one task_span parameter" — the alternative to having both dataset_item and task_outputs. So this is a valid scorer:
def spans_are_present(task_span=None):
return score_result.ScoreResult(name="spans_are_present", value=1.0 if task_span else 0.0)
It never runs. Every score() call ends in:
TypeError: spans_are_present() got an unexpected keyword argument 'dataset_item'
Reproduction, root cause, suggested direction
Reproduction
from opik.evaluation.scorers import scorer_wrapper_metric
def spans_are_present(task_span=None):
return score_result.ScoreResult(name="spans_are_present", value=1.0)
metric = scorer_wrapper_metric.wrap_scorer_functions([spans_are_present], project_name=None)[0]
metric.score(dataset_item={"input": "hi"}, task_outputs={"output": "hello"})
Run against main (2438f12fe2), with only the SDK installed:
wrap_scorer_functions succeeds — the ValueError from validate_scorer_function is not raised.
metric.score(...) raises TypeError: spans_are_present() got an unexpected keyword argument 'dataset_item'.
Why
ScorerWrapperMetric.score has two ways to reach the scorer:
if needs_task_span and task_span is not None:
return self.scorer(dataset_item=..., task_outputs=..., task_span=task_span)
return self.scorer(dataset_item=dataset_item, task_outputs=task_outputs)
The second branch is reached whenever no span is bound, and it always passes dataset_item= and task_outputs=. A scorer that takes neither therefore gets two keywords it does not declare.
The two helpers already draw the distinction correctly — has_task_span_in_parameters is True and requires_task_span_argument is False here, because task_span has a default — but that only routes around the "missing required argument" report; nothing filters the keywords before the call.
Impact
A user whose scorer needs nothing but the span cannot re-score or evaluate at all: every item fails with a TypeError from user code rather than the ScoreMethodMissingArguments the tolerance rules know about, so with a strict error_tolerance the run aborts and with a lenient one every score is marked failed for the wrong reason. The docstring for evaluate_experiment says a scorer "whose task_span parameter has a default still runs", so this contradicts documented behaviour.
Suggested direction
I would filter the call by the scorer's own signature — pass only the arguments it actually declares, so a task-span-only scorer is called as scorer(task_span=None). That keeps the existing behaviour for scorers that do declare dataset_item/task_outputs and fixes the ones that do not. An alternative is to reject task-span-only scorers at construction time, which is simpler but breaks callers who have one today and would then get a ValueError instead of a score.
Happy to open a PR either way — please say which you prefer.
A scorer that declares only
task_spanis accepted at construction and then fails on every call.What happens
validate_scorer_functionaccepts a scorer with "at least onetask_spanparameter" — the alternative to having bothdataset_itemandtask_outputs. So this is a valid scorer:It never runs. Every
score()call ends in:Reproduction, root cause, suggested direction
Reproduction
Run against
main(2438f12fe2), with only the SDK installed:wrap_scorer_functionssucceeds — theValueErrorfromvalidate_scorer_functionis not raised.metric.score(...)raisesTypeError: spans_are_present() got an unexpected keyword argument 'dataset_item'.Why
ScorerWrapperMetric.scorehas two ways to reach the scorer:The second branch is reached whenever no span is bound, and it always passes
dataset_item=andtask_outputs=. A scorer that takes neither therefore gets two keywords it does not declare.The two helpers already draw the distinction correctly —
has_task_span_in_parametersisTrueandrequires_task_span_argumentisFalsehere, becausetask_spanhas a default — but that only routes around the "missing required argument" report; nothing filters the keywords before the call.Impact
A user whose scorer needs nothing but the span cannot re-score or evaluate at all: every item fails with a
TypeErrorfrom user code rather than theScoreMethodMissingArgumentsthe tolerance rules know about, so with a stricterror_tolerancethe run aborts and with a lenient one every score is marked failed for the wrong reason. The docstring forevaluate_experimentsays a scorer "whosetask_spanparameter has a default still runs", so this contradicts documented behaviour.Suggested direction
I would filter the call by the scorer's own signature — pass only the arguments it actually declares, so a task-span-only scorer is called as
scorer(task_span=None). That keeps the existing behaviour for scorers that do declaredataset_item/task_outputsand fixes the ones that do not. An alternative is to reject task-span-only scorers at construction time, which is simpler but breaks callers who have one today and would then get aValueErrorinstead of a score.Happy to open a PR either way — please say which you prefer.