gomoqt ships benchmarks under moqt/*_benchmark_test.go. This document explains
how to run them locally, how CI captures baselines, and how PRs are compared
against those baselines with benchstat.
# 1. Quick sanity check — runs each benchmark once with allocations reporting.
mage bench:short
# 2. Capture a baseline (10 samples, the same shape CI uses).
mage bench:full > bench-baseline.txt
# 3. Make your change, then compare.
mage bench:compare # compares against ./bench-baseline.txt
# or point at a specific baseline:
mage bench:compare path/to/baseline.txtYou can also run the raw commands directly:
# Capture a baseline.
go test -run='^$' -bench=. -benchmem -benchtime=1x -count=10 -timeout=20m ./moqt/... \
| tee bench-baseline.txt
# After your change, capture again...
go test -run='^$' -bench=. -benchmem -benchtime=1x -count=10 -timeout=20m ./moqt/... \
| tee bench-new.txt
# ...and compare.
go install golang.org/x/perf/cmd/benchstat@latest
benchstat bench-baseline.txt bench-new.txtbenchstat needs multiple samples per benchmark to distinguish a real change from
noise. With -count=1 there is no variance estimate and benchstat reports every
delta as significant. -count=10 is the smallest count that gives a usable
variance estimate while keeping the run affordable.
1x runs each benchmark function body exactly once per iteration. The hot loops
are already inside b.N, so 1x exercises the same code paths as the default
1s budget at a fraction of the wall-clock cost. For higher-fidelity numbers on
a fast machine, bump to -benchtime=100ms or -benchtime=1s.
The Go workflow has three jobs relevant to
performance:
| Job | Trigger | Purpose |
|---|---|---|
race |
push to main, PRs touching Go | go test -race ./moqt/... (needs CGO + gcc; ubuntu runners have it) |
benchmark |
push to main, tags v*, manual dispatch |
Captures bench-new.txt and uploads it as the bench-baseline artifact (90-day retention) |
benchmark |
PRs touching Go | Runs benchmarks, downloads the latest main baseline, runs benchstat, uploads a bench-pr-<n> artifact (14-day retention) |
On a PR, the benchmark job:
- Resolves the most recent successful
pushrun ofgo.ymlonmainusinggh run list, then downloads itsbench-baselineartifact. (The artifact is namedbench-new.txtat upload time; it's renamed tobench-baseline.txtlocally.) - Runs
go test -run='^$' -bench=. -benchmem -benchtime=1x -count=10 ./moqt/...to producebench-new.txt. - Runs
benchstat baseline/bench-baseline.txt bench-new.txt. The output table has columns forns/op,B/op,allocs/op, plus variance.~means no significant change;+x%/-x%is a drift beyond noise (p < 0.05).
The comparison step is informational — it continue-on-errors because benchstat
treats improvements as a non-zero exit too. Reviewers should look at the printed
table and the uploaded bench-pr-<n> artifact. If the first PR after a main
merge has no baseline to compare against (no successful main run yet), the step
is skipped gracefully.
benchstat's variance columns matter as much as the mean. A benchmark whose variance jumped several-fold between baseline and PR is unreliable — the "mean drift" it reports is probably noise. When evaluating an optimization PR:
- Look for mean drift (
ns/op,B/op,allocs/op) — the actual signal. - Look for variance inflation — if variance roughly doubled, re-run with a
larger
-benchtimeor higher-countbefore trusting the delta.
go test -race requires CGO and a C compiler. The CI race job sets
CGO_ENABLED=1 and relies on gcc shipped on ubuntu-latest. Local machines
without a C compiler cannot run -race; that is expected. Run it via CI or on a
machine that has gcc/clang installed.