Conversation
--not_fast silently ignored --start_date/--end_date (no date parameters at all) and never forwarded --release_version to load_dataset, so the original-benchmark path loaded the full dataset at the Hub's default revision with no signal. selfrepair dropped the dates; testoutputprediction and codeexecution accepted release_version but never used it. Thread the flags through every loader: not_fast now forwards version_tag and applies the same contest-date window as the fast path, selfrepair passes its dates through, and the test-generation and execution loaders forward version_tag. Fixes LiveCodeBench#153
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
--start_dateand--end_dateselect the contamination window and--release_versionpins the dataset revision, but all three are honored only on the defaultcode_generation_litepath. On--not_fast(the README's documented way to run the original benchmark with the full private test suites),load_code_generation_dataset_not_fasthas no date parameters at all and never forwardsrelease_versiontoload_dataset, so the request silently resolves to whatever the Hub repo's default config currently is and every problem is loaded, including problems outside the requested window.selfrepairdrops the dates;testoutputpredictionandcodeexecutionacceptrelease_versionand never use it. No warning or error mentions any of this, so a run's contamination-free property is decided by the union of loaded problems, not by the flags on the command line.Fixes #153.
Changes
load_code_generation_dataset_not_fastnow takesstart_date/end_date, forwardsversion_tag=release_version, and applies the same contest-date filtering as the fast path.build_prompt_benchmarkpasses the date flags on thenot_fastandselfrepairbranches.load_test_prediction_datasetandload_code_execution_datasetforwardversion_tag=release_version.Verification
The repo has no test suite, so the change ships with a self-contained verifier (https://gist.github.com/AUTHENSOR/6583bcdd3863d7b54b6db3884f9933f1) that drives the real
build_prompt_benchmarkwith the HuggingFaceload_datasetboundary stubbed to record its arguments and return four synthetic problems (two inside 2024-01-01..2024-06-01, one before, one after). Five arms:On current
mainthe same verifier fails at thenot_fastarm exactly as the issue describes: 4/4 problems loaded, out-of-window problems present, andload_datasetcalled with noversion_tag({'split': 'test'}only).