--- id: SMH-PLAN-2_O9OFR0 type: plan title: "Benchmarking: Budgets, A/B, Profiling" spec: SMH-SPEC-SPEC0001 status: approved depends_on: [SMH-PLAN-DFOG846Z] --- # Benchmarking: Budgets, A/B, Profiling ## Entry - RULES `## Testing` names instruction budgets and wall-clock A/B; `## Memory` says the allocator is a benchmark result. - Spec targets: help within 100 ms, frames within 16 ms, encoding 1,000 entries within 5 ms. - Allocation-count bounds exist (`smith-alloc`); one criterion bench exists (`smith/benches/message_creation.rs`). - `target-cpu=native` from the user config is not applied in this repo, so instruction counts depend on the toolchain, not the machine. ## Decisions - Three signals, each where it is strong: allocation counts (exact, tests), instruction counts (`gungraun` over callgrind, exact per toolchain, CI gate), wall clock (`criterion`, `hyperfine`, local A/B, never committed). - Budgets are pinned per bench in `smith-bench/budgets.toml` for one platform (toolchain and libc, the CI image); the gate is a ±2 % band on `Ir` on that platform and a report elsewhere (glibc 2.41 versus 2.44 moved counts by 3 to 23 %); both directions fail so pins stay honest; re-pin is a deliberate commit made inside the CI image (`docker run … rust:latest … cargo x perf --pin`). - One workspace member `smith-bench` owns every bench; it depends on any workspace crate; nothing depends on it; it is a default member so `cargo x check` type-checks benches (`cargo test --benches` is not run). - `gungraun-runner` is a build tool installed by xtask at a pinned version matching the crate. - Wall-clock spec targets are informational rows in the report bundle, never a gate. ## Order 1. `smith-bench` crate: library benches for frame encode and decode of 1,000 entries, `Session::from_frames`, `fork_at`, agent request build with masking, one SSE fixture per adapter, `canonical_json`, `handle_line`; binary bench for `smith --help` and `smith eval --mock`; criterion bench moved from `smith`. 2. `budgets.toml`: one `ir` pin per bench id, written by `cargo x perf --pin`. 3. `cargo x perf`: install runner, run gungraun benches with `--save-summary=json`, compare every `Ir` to its pin within ±2 %, write `target/report/perf.md`; `--pin` rewrites budgets; `cargo x bench` runs criterion with passthrough. 4. CI: `.ci/perf.kdl` per push (valgrind installed in the job); `report` bundle gains the perf section and a hyperfine row for `smith --help` and `smith eval --mock`. 5. Research note and TO_DISCUSS cleanup; deslop and compress passes. ## Interfaces - `xtask`: `perf [--pin]`, `bench [args]`; `bootstrap` installs `gungraun-runner`. - `smith-bench/budgets.toml`: `[::::::] ir = N`. - `.ci/perf.kdl`. ## Risks - Callgrind runs are slow; the set stays small and every bench runs one shot. - Toolchain bumps shift counts; the re-pin commit names the toolchain. - Binary benches include process start-up; that is the point for `--help`, noise for anything else. ## Exit - `cargo x perf` green with pins; CI `perf` job observed; report bundle carries the perf section.