--- name: benchmarking description: "Use before claiming or chasing a performance difference, choosing between two implementations, layouts, or allocators, or pinning a performance budget: pick the signal that answers the question, compare fairly, and keep only what is reproducible." --- # Benchmarking Measure the question, not the machine. ## Pick the signal | question | signal | tool | |---|---|---| | does this path allocate more | allocation count, bytes | counting allocator in a test | | did this change do more work | instructions, est. cycles, cache hits | callgrind via `gungraun`, one shot, deterministic | | is A faster than B on this box | wall clock | `criterion` (lib), `hyperfine` (binary), same machine, same session | | where does the time go | profile | `perf stat -d`, `perf record` + flamegraph, `callgrind_annotate` | | is the type or record too big | size | `size_of` assertion | Wall clock answers "faster here, now"; instructions answer "less work"; only the second travels between machines. ## Compare fairly - A/B on one machine, one session, back to back; interleave runs; nothing else running. - Same toolchain, same profile, same inputs; name baselines by change id (`--save-baseline `), never `before`. - Warm up; report median and spread; a difference inside the spread is no difference. - Inputs at the sizes the spec names, and one size larger than fits in cache. - `black_box` inputs and outputs; check the loop was not optimized away. ## Keep only what reproduces - Commit: instruction budgets with a tolerance band, allocation counts, size assertions. - Never commit wall-clock numbers; put them in a report or a commit message with the machine named. - A budget moves only by a deliberate commit that says why (toolchain, layout, algorithm). - An improvement beyond the band fails too until re-pinned; a pin that drifts down silently was never a pin. ## Read the result - Fewer instructions but more time: memory-bound; look at cache misses and layout before code. - Fewer cache misses but more time: unpacking cost ate the win; drop the encoding. - Callgrind counts exclude wall-clock effects (I/O waits, contention, frequency scaling); pair with one wall-clock run before believing a win. - Binary benches include process start-up; that is the measurement for `--help`, noise for everything else. ## Do not - Benchmark in CI for wall clock; CI containers share cores and vary by an order of magnitude. - Optimize before the profile names the hot path. - Add a benchmark without the spec sentence or bound it proves. ## Sources - [gungraun guide](https://gungraun.github.io/gungraun/) - [criterion user guide](https://bheisler.github.io/criterion.rs/book/) - [The Rust Performance Book: benchmarking and profiling](https://nnethercote.github.io/perf-book/benchmarking.html) - [hyperfine](https://github.com/sharkdp/hyperfine)