name: benchmarking
description: "Use before claiming or chasing a performance difference, choosing between two implementations, layouts, or allocators, or pinning a performance budget: pick the signal that answers the question, compare fairly, and keep only what is reproducible."
Benchmarking
Measure the question, not the machine.
Pick the signal
question
signal
tool
does this path allocate more
allocation count, bytes
counting allocator in a test
did this change do more work
instructions, est. cycles, cache hits
callgrind via gungraun, one shot, deterministic
is A faster than B on this box
wall clock
criterion (lib), hyperfine (binary), same machine, same session
where does the time go
profile
perf stat -d, perf record + flamegraph, callgrind_annotate
is the type or record too big
size
size_of assertion
Wall clock answers "faster here, now"; instructions answer "less work"; only the second travels between machines.
Compare fairly
A/B on one machine, one session, back to back; interleave runs; nothing else running.
Same toolchain, same profile, same inputs; name baselines by change id (--save-baseline <change>), never before.
Warm up; report median and spread; a difference inside the spread is no difference.
Inputs at the sizes the spec names, and one size larger than fits in cache.
black_box inputs and outputs; check the loop was not optimized away.
Keep only what reproduces
Commit: instruction budgets with a tolerance band, allocation counts, size assertions.
Never commit wall-clock numbers; put them in a report or a commit message with the machine named.
A budget moves only by a deliberate commit that says why (toolchain, layout, algorithm).
An improvement beyond the band fails too until re-pinned; a pin that drifts down silently was never a pin.
Read the result
Fewer instructions but more time: memory-bound; look at cache misses and layout before code.
Fewer cache misses but more time: unpacking cost ate the win; drop the encoding.
Callgrind counts exclude wall-clock effects (I/O waits, contention, frequency scaling); pair with one wall-clock run before believing a win.
Binary benches include process start-up; that is the measurement for --help, noise for everything else.
Do not
Benchmark in CI for wall clock; CI containers share cores and vary by an order of magnitude.
Optimize before the profile names the hot path.
Add a benchmark without the spec sentence or bound it proves.
---
name: benchmarking
description: "Use before claiming or chasing a performance difference, choosing between two implementations, layouts, or allocators, or pinning a performance budget: pick the signal that answers the question, compare fairly, and keep only what is reproducible."
---
# Benchmarking
Measure the question, not the machine.
## Pick the signal
| question | signal | tool |
|---|---|---|
| does this path allocate more | allocation count, bytes | counting allocator in a test |
| did this change do more work | instructions, est. cycles, cache hits | callgrind via `gungraun`, one shot, deterministic |
| is A faster than B on this box | wall clock | `criterion` (lib), `hyperfine` (binary), same machine, same session |
| where does the time go | profile | `perf stat -d`, `perf record` + flamegraph, `callgrind_annotate` |
| is the type or record too big | size | `size_of` assertion |
Wall clock answers "faster here, now"; instructions answer "less work"; only the second travels between machines.
## Compare fairly
- A/B on one machine, one session, back to back; interleave runs; nothing else running.
- Same toolchain, same profile, same inputs; name baselines by change id (`--save-baseline <change>`), never `before`.
- Warm up; report median and spread; a difference inside the spread is no difference.
- Inputs at the sizes the spec names, and one size larger than fits in cache.
- `black_box` inputs and outputs; check the loop was not optimized away.
## Keep only what reproduces
- Commit: instruction budgets with a tolerance band, allocation counts, size assertions.
- Never commit wall-clock numbers; put them in a report or a commit message with the machine named.
- A budget moves only by a deliberate commit that says why (toolchain, layout, algorithm).
- An improvement beyond the band fails too until re-pinned; a pin that drifts down silently was never a pin.
## Read the result
- Fewer instructions but more time: memory-bound; look at cache misses and layout before code.
- Fewer cache misses but more time: unpacking cost ate the win; drop the encoding.
- Callgrind counts exclude wall-clock effects (I/O waits, contention, frequency scaling); pair with one wall-clock run before believing a win.
- Binary benches include process start-up; that is the measurement for `--help`, noise for everything else.
## Do not
- Benchmark in CI for wall clock; CI containers share cores and vary by an order of magnitude.
- Optimize before the profile names the hot path.
- Add a benchmark without the spec sentence or bound it proves.
## Sources
- [gungraun guide](https://gungraun.github.io/gungraun/)
- [criterion user guide](https://bheisler.github.io/criterion.rs/book/)
- [The Rust Performance Book: benchmarking and profiling](https://nnethercote.github.io/perf-book/benchmarking.html)
- [hyperfine](https://github.com/sharkdp/hyperfine)