Repository navigation
feat: CodeRAG-Bench benchmark Version 0.1 - #2362
boerz-coding wants to merge 39 commits into
Conversation
Done: - Implemented basic logic for downloading and preprocessing the "humaneval" sub-task - Created initial directory structure for CodeRAG benchmark; due to its complexity, it is placed in a separate folder - Added function signatures for the benchmark class Ongoing: - Canonical retrieval implementation for the "humaneval" sub-task To be done: - Support for the remaining 6 sub-tasks - Generation support - Open-retrieval support - test/examples
Completed basic logics for HumanEval benchmark
Temporary draft, need further edit.
- Edited the way retrieval result is saved - Edited the way for generating prompts - Implemented code post-processing Next Step: Implement the code evaluation logic (To evaluate code, we need to actually run it, so is somewhat challenging)
- Implemented evaluation logic for Humaneval, using relevant files from coderag-bench (now in camel/benchmarks/coderag_bench/code_generation_evaluation) - Adjusted the output structure of generation step to match the expected format in compute_code_eval.
- Completed subset logics - Still need further test, currently has dependency issues
- Humaneval is fully implemented and has passed test with a subset of 5 lines of data - Output logs should be cleaned before making PR public. The retriever part need to be considered carefully
Editing docstrings to complete a MVP version that supports full RAG (canonical) + evaluation process for HumanEval subtask.
…/camel into feat-code_rag_bench
Thank you Wendong. This is very helpful Co-authored-by: Wendong-Fan <133094783+Wendong-Fan@users.noreply.github.com>
…/camel into feat-code_rag_bench
…ode_eval and execute
… into feat-code_rag_bench
|
Hi @zjrwtx I have edited all docstrings. Thank you for your helpful comments. |
|
Hi @Wendong-Fan, I’ve addressed all comments from you and Yifeng — the PR is ready for review again. Since this PR was moved from my forked repo to the main repo and your comments was left on the original PR , I wanted to recap the key changes from previous reviews:
|
|
Important Review skippedAuto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the You can disable this status message by setting the Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. 🪧 TipsChatThere are 3 ways to chat with CodeRabbit:
SupportNeed help? Create a ticket on our support page for assistance with any issues or questions. Note: Be mindful of the bot's finite context window. It's strongly recommended to break down tasks such as reading entire modules into smaller chunks. For a focused discussion, use review comments to chat about specific files and their changes, instead of using the PR comments. CodeRabbit Commands (Invoked using PR comments)
Other keywords and placeholders
CodeRabbit Configuration File (
|
Description
I have moved my PR #2199 from my forked repo to the main repo, to allow proper chaining with follow-up PRs and CI testing.
This PR implements Version 0.1 (MVP) of the CodeRAG-Bench benchmark (issue #1462).
Introduction
CodeRAG-Bench is designed to enable rigorous evaluations and advance research on retrieval-augmented code generation.
Reference: CodeRAG-Bench Homepage
Current Status (v0.1)
This MVP supports:
Note:
The full benchmark includes 7 code generation tasks (HumanEval, MBPP, LiveCodeBench, DS-1000, ODEX, RepoEval, SWE-bench-Lite), and both canonical and open-corpus retrieval. This MVP focuses on the first task (HumanEval) with canonical retrieval only.
The goal of this MVP is to provide an early working version for feedback and for potential users to start experimenting.
Example Output Files
Benchmark output files (retrieval results, generations, evaluation metrics) are available here for reproducibility.
Next Steps (Mini Roadmap)
I will continue working towards the full CodeRAG-Bench implementation. I also plan to split tasks into sub-issues to invite community contributions whenever possible.
Planned next steps:
Checklist
Go over all the following points, and put an
xin all the boxes that apply.Fixes #issue-numberin the PR description (required)pyproject.tomlanduv lockIf you are unsure about any of these, don't hesitate to ask. We are here to help!