lcb_runner/evaluation/compute_code_execution_metrics.py:12-19
for g in gs:
if i in g:
pass
else:
code_to_execute = f"{BASE_IMPORTS}\n{c}\nassert {o} == {g}"
execution_results.append(check_correctness(code_to_execute, 3))
if len(execution_results) == 0:
execution_results = [False] * len(gs)
The i in g guard disqualifies a predicted output that contains the function call, which would otherwise pass assert <literal> == f(...) by running the code. Instead of appending False, the loop appends nothing, so the generation leaves the denominator. code_execution_metrics (lines 46-47) then uses n = len(execution_result), and CodeExecutionProblem.insert_output_evaluation uses graded_list.count(True) / len(graded_list).
What happens (reproduced on the real module)
Code execution: generations disqualified by the i in g guard are dropped from pass@1, not scored as False (lcb_runner/evaluation/compute_code_execution_metrics.py:11-18).
evaluate_score skips any extracted answer that contains the input call string. The prompt forbids these answers ("a literal ... no function calls"). The skip appends nothing, and the [False] * len(gs) fallback only fires when every generation was skipped. When only some are skipped, the result list is shorter than gs. That shorter list feeds three places, all with the smaller n:
code_execution_metrics (n = len(execution_result), lines 46-47)
CodeExecutionProblem.insert_output_evaluation (code_execution.py:48)
compute_scores.py totals (lines 95-103)
Reproduced with the real modules: the real extraction, the real check_correctness sandbox and the real insert_output_evaluation. With 5 samples of the prompt's performOperation example:
- 1 correct literal + 4 answers
assert performOperation(s = "hi") == performOperation(s = "hi"): graded_list == [True], pass@1 = 1.0. The same run with 4 wrong literals gives 0.2.
- 1 wrong + 4 echoes: pass@1 is 0.0, but graded_list has length 1 against a pred_list of 5, so saved verdicts can no longer be matched to generations by position.
The logic comes from CRUXEval (utils_general.py, fallback if True not in execution_results). LCB already changed the fallback to len == 0. The all-disqualified fallback shows the intent is to score disqualified answers as False.
Fix: execution_results.append(False) in the i in g branch, drop the now-redundant fallback, and add assert len(execution_results) == len(gs). This makes the code match the prompt's rules and the codegen path, which keeps one entry per generation.
How often models echo the exact call has not been measured, so the change in leaderboard numbers is expected to be small. Note that i in g is an exact substring match: a call written with different spacing or argument style is not caught. That is a separate gap this fix does not close.
Happy to open the PR.
lcb_runner/evaluation/compute_code_execution_metrics.py:12-19The
i in gguard disqualifies a predicted output that contains the function call, which would otherwise passassert <literal> == f(...)by running the code. Instead of appendingFalse, the loop appends nothing, so the generation leaves the denominator.code_execution_metrics(lines 46-47) then usesn = len(execution_result), andCodeExecutionProblem.insert_output_evaluationusesgraded_list.count(True) / len(graded_list).What happens (reproduced on the real module)
Code execution: generations disqualified by the
i in gguard are dropped from pass@1, not scored as False (lcb_runner/evaluation/compute_code_execution_metrics.py:11-18).evaluate_scoreskips any extracted answer that contains the input call string. The prompt forbids these answers ("a literal ... no function calls"). The skip appends nothing, and the[False] * len(gs)fallback only fires when every generation was skipped. When only some are skipped, the result list is shorter thangs. That shorter list feeds three places, all with the smaller n:code_execution_metrics(n = len(execution_result), lines 46-47)CodeExecutionProblem.insert_output_evaluation(code_execution.py:48)compute_scores.pytotals (lines 95-103)Reproduced with the real modules: the real extraction, the real
check_correctnesssandbox and the realinsert_output_evaluation. With 5 samples of the prompt'sperformOperationexample:assert performOperation(s = "hi") == performOperation(s = "hi"): graded_list == [True], pass@1 = 1.0. The same run with 4 wrong literals gives 0.2.The logic comes from CRUXEval (utils_general.py, fallback
if True not in execution_results). LCB already changed the fallback tolen == 0. The all-disqualified fallback shows the intent is to score disqualified answers as False.Fix:
execution_results.append(False)in thei in gbranch, drop the now-redundant fallback, and addassert len(execution_results) == len(gs). This makes the code match the prompt's rules and the codegen path, which keeps one entry per generation.How often models echo the exact call has not been measured, so the change in leaderboard numbers is expected to be small. Note that
i in gis an exact substring match: a call written with different spacing or argument style is not caught. That is a separate gap this fix does not close.Happy to open the PR.