Verbosity Measurement
Verbosity quantifies the sheer volume of code each model generates to solve identical tasks.
- Lines of Code (LOC)
Total number of lines of code generated across all 4,442 tasks, including blank lines and comments. This metric reveals whether a model tends toward concise or elaborate implementations. - Token Count
Total tokens generated in the code output, providing a language-agnostic measure of code volume that accounts for the actual content density. - Code Density
Ratio of executable statements to total lines, indicating how compact or spread out the code structure is.
Complexity Measurement
Complexity metrics evaluate the structural and logical intricacy of the generated code using industry-standard software engineering measures.
- Cyclomatic Complexity
Measures the number of linearly independent paths through the code. Higher values indicate more complex control flow with multiple branches, loops, and decision points. - Cognitive Complexity
Evaluates how difficult code is for humans to understand, accounting for nested structures and non-linear control flow. This metric better reflects maintainability concerns than cyclomatic complexity alone. - Nesting Depth
Maximum depth of nested structures (loops, conditionals, try-catch blocks). Deeper nesting typically correlates with harder-to-maintain code.
Communication & Documentation
Documentation metrics measure how thoroughly each model explains its code through comments and inline documentation.
- Comment Density (%)
Percentage of lines containing comments relative to total lines of code. Calculated as (comment lines / total lines) × 100. - Docstring Coverage
Percentage of functions, classes, and modules that include docstrings or header comments explaining their purpose and usage. - Inline Explanation Quality
Qualitative assessment of whether comments add meaningful context or merely restate obvious code operations.
Software Quality Analysis
Quality metrics identify issues across three critical dimensions using static analysis tools.
- Reliability (Bugs)
Code defects that would cause runtime errors, incorrect behavior, or system failures. Includes control-flow mistakes, resource leaks, and concurrency issues. - Security (Vulnerabilities)
Weaknesses that could be exploited by attackers, such as injection flaws, path traversal vulnerabilities, and insecure cryptographic practices. Severity ranges from INFO to BLOCKER. - Maintainability (Code Smells)
Design weaknesses that increase technical debt and make code harder to modify. Includes duplicated code, overly complex functions, and poor naming conventions.
Functional Performance
Pass rate measures each model's ability to produce functionally correct solutions.
- Pass Rate (%)
Percentage of the 4,442 tasks where the generated code produces correct output for all test cases. This establishes the baseline capability before considering code quality. - Issue Density
Total number of reliability, security, and maintainability issues per thousand lines of code (issues/KLOC). Lower values indicate cleaner code.


