Add a scan benchmark suite and a CI performance gate - #485
Conversation
New crate crates/socket-patch-bench (publish = false, no new third-party dependencies) that benchmarks `socket-patch scan` end to end: - Synthetic, seed-deterministic projects for every package manager in its native lockfile format and install layout: npm, pnpm (isolated store), yarn classic, yarn berry, bun, vlt, pip requirements, uv, PEP 751 pylock, poetry, pipenv, pdm, bundler, composer, cargo, go, nuget and maven. Per-user caches live under the fixture's own HOME. - A std-only mock of the patch API, proxy and artifact host that answers from the scenario's catalog and counts requests per endpoint. - Each PM runs a hosted scan of a fresh project and a rescan of an already-redirected one; npm also runs --dry-run, the public proxy and 40 ms simulated latency (request concurrency). - Every run is validated (scanned/lockfile-only/patched counts, redirected patches, exact rewritten files that really changed on disk, allowed warnings, no unexpected requests, dry runs write nothing), so a scan that skips work fails instead of looking fast. Runs start from the pristine tree via snapshot diff/restore, with a cleared env. - `compare` interleaves base and head runs on one machine and fails on a wall/CPU slowdown (median paired ratio past 10% with a 95% sign-test interval above 1, confirmed by a second round), peak RSS growth, more API requests, or a head that fails validation. .github/workflows/bench.yml builds the PR's base and head with the new `perf` profile (release codegen, thin LTO) and runs the comparison on every PR touching Rust code; pushes to main are recorded without gating. The `performance-regression-accepted` label turns a regression into a report. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017hKPGJkJfyhsQrzXYf7RJZ
Intra-store dependency symlinks live at {store}/{key}/node_modules/{dep},
so only the dependency name's scope adds a directory level. Counting the
parent's slashes pointed every link under a scoped parent one directory
too high, leaving those packages with a broken isolated layout.
Also replace the bare unwraps on the per-binary result maps with expect()
messages stating the invariant, per Bugbot review.
Co-Authored-By: Claude <noreply@anthropic.com>
|
bugbot run Generated by Claude Code |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit 08936bd. Configure here.
|
Burn-down agent: ready for review at
Slack announcement not sent: no Slack send tool is available to this agent; the next run will retry. Generated by Claude Code |
|
Reviewed No blocking findings in the benchmark harness, fixture restoration, paired comparison/gating, or CI workflow. Checked baseline selection, the regression-acceptance label, and validation-failure handling. Validation: Ran |
|
Second-pass integration check of On the combined tree, 29 benchmark harness tests passed, and I rebuilt both the CLI and harness, then ran all 39 scenarios with This verifies functional compatibility only; the development-profile smoke does not establish timing or memory regression results. |
Summary
Adds
crates/socket-patch-bench(unpublished, no new third-party crates inCargo.lock) and.github/workflows/bench.yml, a CI gate that compares every PR'ssocket-patch scanperformance against its base.What is measured
The harness generates a seed-deterministic project per package manager in its native lockfile format and install layout, then runs the real binary against a local, std-only mock of the patch API, the public proxy and the artifact host. Package managers: npm, pnpm (isolated store), yarn classic, yarn berry, bun, vlt, pip, uv, PEP 751 pylock, poetry, pipenv, pdm, bundler, composer, cargo, go, nuget, maven.
<pm>/hosted: the defaultscanon a freshly installed project (crawl, lockfile inventory, batch query, details, references, artifact checks, rewrite, views, writes).<pm>/rescan: the same scan on an already-redirected project (pin discovery and update detection over rewritten locks).--dry-run, the public proxy (no token), and 40 ms simulated latency.Why the numbers can be trusted
env -ienvironment.HOMEand the caches point into the fixture, telemetry and the update check are off, and proxies point at a closed port. strace shows the CLI spawns nothing.The gate
The workflow builds base and head with the new
perfprofile (release codegen, thin LTO) on one runner and interleaves their runs. A PR fails when any of these hold:The
performance-regression-acceptedlabel turns a regression into a report instead of a failure. Pushes tomainare recorded as artifacts without gating.Validation
--ecosystems pypion an npm project) was rejected as invalid.cargo clippy --workspace --all-features -D warningspasses, and the crate's 29 unit tests pass.Notes
scan --mode vendored/agent, which need real patched archives, and Deno, which has no hosted rewrite.gopatchmodules ingo.sumas extra lockfile-only packages (412 vs 400 scanned). That looks like a double count; it is encoded as current behavior and not changed here.🤖 Generated with Claude Code
https://claude.ai/code/session_017hKPGJkJfyhsQrzXYf7RJZ
Generated by Claude Code
Note
Low Risk
Adds CI tooling and an unpublished benchmark crate; it does not change shipped scan behavior, though flaky gates could block PRs until labeled or tuned.
Overview
Introduces a
scanperformance gate for Rust changes: a new unpublished workspace cratesocket-patch-benchplus.github/workflows/bench.yml.The harness builds deterministic synthetic projects per package manager (npm family, Python tools, Ruby, PHP, Rust, Go, NuGet, Maven), serves patches from a local HTTP mock, and runs the real
socket-patch scanunder an isolated environment. Each sample is validated (JSON counts, redirects, rewrites, request patterns) so skipped work cannot look like a win. The CLI supportslist,run,compare, andserve(profiling), emitting markdown andresults.json.CI builds base and head with a new
[profile.perf](release semantics, thin LTO), runs interleaved paired comparisons, and fails PRs on statistically significant wall/CPU slowdown, RSS growth, extra API calls, or head validation failures.performance-regression-acceptedsuppresses timing/memory/request failures only;mainpushes record results without gating. Docs now point performance work at this suite; record/replay remains for real API traffic.Reviewed by Cursor Bugbot for commit 08936bd. Configure here.
Generated by Claude Code