Luigit
repositories / smith

smith

There are many coding harnesses - but this one is fast

owned by admin

.pi/skills/benchmarking/SKILL.md

Raw
Rendered preview

name: benchmarking description: "Use before claiming or chasing a performance difference, choosing between two implementations, layouts, or allocators, or pinning a performance budget: pick the signal that answers the question, compare fairly, and keep only what is reproducible."

Benchmarking

Measure the question, not the machine.

Pick the signal

question signal tool
does this path allocate more allocation count, bytes counting allocator in a test
did this change do more work instructions, est. cycles, cache hits callgrind via gungraun, one shot, deterministic
is A faster than B on this box wall clock criterion (lib), hyperfine (binary), same machine, same session
where does the time go profile perf stat -d, perf record + flamegraph, callgrind_annotate
is the type or record too big size size_of assertion

Wall clock answers "faster here, now"; instructions answer "less work"; only the second travels between machines.

Compare fairly

  • A/B on one machine, one session, back to back; interleave runs; nothing else running.
  • Same toolchain, same profile, same inputs; name baselines by change id (--save-baseline <change>), never before.
  • Warm up; report median and spread; a difference inside the spread is no difference.
  • Inputs at the sizes the spec names, and one size larger than fits in cache.
  • black_box inputs and outputs; check the loop was not optimized away.

Keep only what reproduces

  • Commit: instruction budgets with a tolerance band, allocation counts, size assertions.
  • Never commit wall-clock numbers; put them in a report or a commit message with the machine named.
  • A budget moves only by a deliberate commit that says why (toolchain, layout, algorithm).
  • An improvement beyond the band fails too until re-pinned; a pin that drifts down silently was never a pin.

Read the result

  • Fewer instructions but more time: memory-bound; look at cache misses and layout before code.
  • Fewer cache misses but more time: unpacking cost ate the win; drop the encoding.
  • Callgrind counts exclude wall-clock effects (I/O waits, contention, frequency scaling); pair with one wall-clock run before believing a win.
  • Binary benches include process start-up; that is the measurement for --help, noise for everything else.

Do not

  • Benchmark in CI for wall clock; CI containers share cores and vary by an order of magnitude.
  • Optimize before the profile names the hot path.
  • Add a benchmark without the spec sentence or bound it proves.

Sources

---
name: benchmarking
description: "Use before claiming or chasing a performance difference, choosing between two implementations, layouts, or allocators, or pinning a performance budget: pick the signal that answers the question, compare fairly, and keep only what is reproducible."
---

# Benchmarking

Measure the question, not the machine.

## Pick the signal

| question | signal | tool |
|---|---|---|
| does this path allocate more | allocation count, bytes | counting allocator in a test |
| did this change do more work | instructions, est. cycles, cache hits | callgrind via `gungraun`, one shot, deterministic |
| is A faster than B on this box | wall clock | `criterion` (lib), `hyperfine` (binary), same machine, same session |
| where does the time go | profile | `perf stat -d`, `perf record` + flamegraph, `callgrind_annotate` |
| is the type or record too big | size | `size_of` assertion |

Wall clock answers "faster here, now"; instructions answer "less work"; only the second travels between machines.

## Compare fairly

- A/B on one machine, one session, back to back; interleave runs; nothing else running.
- Same toolchain, same profile, same inputs; name baselines by change id (`--save-baseline <change>`), never `before`.
- Warm up; report median and spread; a difference inside the spread is no difference.
- Inputs at the sizes the spec names, and one size larger than fits in cache.
- `black_box` inputs and outputs; check the loop was not optimized away.

## Keep only what reproduces

- Commit: instruction budgets with a tolerance band, allocation counts, size assertions.
- Never commit wall-clock numbers; put them in a report or a commit message with the machine named.
- A budget moves only by a deliberate commit that says why (toolchain, layout, algorithm).
- An improvement beyond the band fails too until re-pinned; a pin that drifts down silently was never a pin.

## Read the result

- Fewer instructions but more time: memory-bound; look at cache misses and layout before code.
- Fewer cache misses but more time: unpacking cost ate the win; drop the encoding.
- Callgrind counts exclude wall-clock effects (I/O waits, contention, frequency scaling); pair with one wall-clock run before believing a win.
- Binary benches include process start-up; that is the measurement for `--help`, noise for everything else.

## Do not

- Benchmark in CI for wall clock; CI containers share cores and vary by an order of magnitude.
- Optimize before the profile names the hot path.
- Add a benchmark without the spec sentence or bound it proves.

## Sources

- [gungraun guide](https://gungraun.github.io/gungraun/)
- [criterion user guide](https://bheisler.github.io/criterion.rs/book/)
- [The Rust Performance Book: benchmarking and profiling](https://nnethercote.github.io/perf-book/benchmarking.html)
- [hyperfine](https://github.com/sharkdp/hyperfine)