Skip to content

Code execution: generations disqualified by the input-echo guard are dropped from the pass@1 denominator instead of being scored False #158

Description

@shaurya416

lcb_runner/evaluation/compute_code_execution_metrics.py:12-19

for g in gs:
    if i in g:
        pass
    else:
        code_to_execute = f"{BASE_IMPORTS}\n{c}\nassert {o} == {g}"
        execution_results.append(check_correctness(code_to_execute, 3))
if len(execution_results) == 0:
    execution_results = [False] * len(gs)

The i in g guard disqualifies a predicted output that contains the function call, which would otherwise pass assert <literal> == f(...) by running the code. Instead of appending False, the loop appends nothing, so the generation leaves the denominator. code_execution_metrics (lines 46-47) then uses n = len(execution_result), and CodeExecutionProblem.insert_output_evaluation uses graded_list.count(True) / len(graded_list).

What happens (reproduced on the real module)

Code execution: generations disqualified by the i in g guard are dropped from pass@1, not scored as False (lcb_runner/evaluation/compute_code_execution_metrics.py:11-18).

evaluate_score skips any extracted answer that contains the input call string. The prompt forbids these answers ("a literal ... no function calls"). The skip appends nothing, and the [False] * len(gs) fallback only fires when every generation was skipped. When only some are skipped, the result list is shorter than gs. That shorter list feeds three places, all with the smaller n:

  • code_execution_metrics (n = len(execution_result), lines 46-47)
  • CodeExecutionProblem.insert_output_evaluation (code_execution.py:48)
  • compute_scores.py totals (lines 95-103)

Reproduced with the real modules: the real extraction, the real check_correctness sandbox and the real insert_output_evaluation. With 5 samples of the prompt's performOperation example:

  • 1 correct literal + 4 answers assert performOperation(s = "hi") == performOperation(s = "hi"): graded_list == [True], pass@1 = 1.0. The same run with 4 wrong literals gives 0.2.
  • 1 wrong + 4 echoes: pass@1 is 0.0, but graded_list has length 1 against a pred_list of 5, so saved verdicts can no longer be matched to generations by position.

The logic comes from CRUXEval (utils_general.py, fallback if True not in execution_results). LCB already changed the fallback to len == 0. The all-disqualified fallback shows the intent is to score disqualified answers as False.

Fix: execution_results.append(False) in the i in g branch, drop the now-redundant fallback, and add assert len(execution_results) == len(gs). This makes the code match the prompt's rules and the codegen path, which keeps one entry per generation.

How often models echo the exact call has not been measured, so the change in leaderboard numbers is expected to be small. Note that i in g is an exact substring match: a call written with different spacing or argument style is not caught. That is a separate gap this fix does not close.

Happy to open the PR.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions