Celery Organization reposted this
Introducing FrontierCode: a new coding eval by Cognition that raises the bar for difficulty and quality. Most coding benchmarks ask: does the code pass unit tests? FrontierCode asks a harder question: would you actually merge it? Models can increasingly write code that works, but is sloppy, brittle, or hard to maintain. FrontierCode is designed to measure whether AI-generated code meets the standards of real-world software engineering. Built with leading open-source maintainers, FrontierCode includes realistic coding tasks from 36 flagship repositories, including Celery and Budibase. More than 20 world-class developers contributed tasks on repos they maintain, with each task taking 40+ hours of work and multiple rounds of iteration. To evaluate mergeability, FrontierCode combines custom rubrics, novel verifiers, and tests to check for correctness, test quality, scope discipline, style, and adherence to codebase standards. Every task in the dataset goes through an extensive QC pipeline with adversarial testing, calibration, multi-stage review, and manual review by a Cognition researcher. This reduces both false positives and false negatives, producing 81% fewer misclassification errors than SWE-Bench Pro. The benchmark is intentionally difficult and diverse. Tasks have concise problem statements, but require large solutions across multiple files, languages, and problem types. FrontierCode includes three task sets: 1. Extended: 150 tasks 2. Main: 100 tasks 3. Diamond: 50 tasks State-of-the-art LLMs still have significant room to improve. The top model scores just 13.4/100 on the Diamond task set. We built FrontierCode to move software engineering evals closer to the standard that matters in production: code that maintainers would trust and merge. Read the full model results and technical implementation details: https://lnkd.in/eTQufgGi