Luigit
repositories / smith

smith

There are many coding harnesses - but this one is fast

owned by admin

.system/plans/SMH-PLAN-2_O9OFR0-benchmarking/index.md

Raw
Rendered preview

id: SMH-PLAN-2_O9OFR0 type: plan title: "Benchmarking: Budgets, A/B, Profiling" spec: SMH-SPEC-SPEC0001 status: approved depends_on: [SMH-PLAN-DFOG846Z]

Benchmarking: Budgets, A/B, Profiling

Entry

  • RULES ## Testing names instruction budgets and wall-clock A/B; ## Memory says the allocator is a benchmark result.
  • Spec targets: help within 100 ms, frames within 16 ms, encoding 1,000 entries within 5 ms.
  • Allocation-count bounds exist (smith-alloc); one criterion bench exists (smith/benches/message_creation.rs).
  • target-cpu=native from the user config is not applied in this repo, so instruction counts depend on the toolchain, not the machine.

Decisions

  • Three signals, each where it is strong: allocation counts (exact, tests), instruction counts (gungraun over callgrind, exact per toolchain, CI gate), wall clock (criterion, hyperfine, local A/B, never committed).
  • Budgets are pinned per bench in smith-bench/budgets.toml for one platform (toolchain and libc, the CI image); the gate is a ±2 % band on Ir on that platform and a report elsewhere (glibc 2.41 versus 2.44 moved counts by 3 to 23 %); both directions fail so pins stay honest; re-pin is a deliberate commit made inside the CI image (docker run … rust:latest … cargo x perf --pin).
  • One workspace member smith-bench owns every bench; it depends on any workspace crate; nothing depends on it; it is a default member so cargo x check type-checks benches (cargo test --benches is not run).
  • gungraun-runner is a build tool installed by xtask at a pinned version matching the crate.
  • Wall-clock spec targets are informational rows in the report bundle, never a gate.

Order

  1. smith-bench crate: library benches for frame encode and decode of 1,000 entries, Session::from_frames, fork_at, agent request build with masking, one SSE fixture per adapter, canonical_json, handle_line; binary bench for smith --help and smith eval --mock; criterion bench moved from smith.
  2. budgets.toml: one ir pin per bench id, written by cargo x perf --pin.
  3. cargo x perf: install runner, run gungraun benches with --save-summary=json, compare every Ir to its pin within ±2 %, write target/report/perf.md; --pin rewrites budgets; cargo x bench runs criterion with passthrough.
  4. CI: .ci/perf.kdl per push (valgrind installed in the job); report bundle gains the perf section and a hyperfine row for smith --help and smith eval --mock.
  5. Research note and TO_DISCUSS cleanup; deslop and compress passes.

Interfaces

  • xtask: perf [--pin], bench [args]; bootstrap installs gungraun-runner.
  • smith-bench/budgets.toml: [<file>::<group>::<function>::<id>] ir = N.
  • .ci/perf.kdl.

Risks

  • Callgrind runs are slow; the set stays small and every bench runs one shot.
  • Toolchain bumps shift counts; the re-pin commit names the toolchain.
  • Binary benches include process start-up; that is the point for --help, noise for anything else.

Exit

  • cargo x perf green with pins; CI perf job observed; report bundle carries the perf section.
---
id: SMH-PLAN-2_O9OFR0
type: plan
title: "Benchmarking: Budgets, A/B, Profiling"
spec: SMH-SPEC-SPEC0001
status: approved
depends_on: [SMH-PLAN-DFOG846Z]
---

# Benchmarking: Budgets, A/B, Profiling

## Entry

- RULES `## Testing` names instruction budgets and wall-clock A/B; `## Memory` says the allocator is a benchmark result.
- Spec targets: help within 100 ms, frames within 16 ms, encoding 1,000 entries within 5 ms.
- Allocation-count bounds exist (`smith-alloc`); one criterion bench exists (`smith/benches/message_creation.rs`).
- `target-cpu=native` from the user config is not applied in this repo, so instruction counts depend on the toolchain, not the machine.

## Decisions

- Three signals, each where it is strong: allocation counts (exact, tests), instruction counts (`gungraun` over callgrind, exact per toolchain, CI gate), wall clock (`criterion`, `hyperfine`, local A/B, never committed).
- Budgets are pinned per bench in `smith-bench/budgets.toml` for one platform (toolchain and libc, the CI image); the gate is a ±2 % band on `Ir` on that platform and a report elsewhere (glibc 2.41 versus 2.44 moved counts by 3 to 23 %); both directions fail so pins stay honest; re-pin is a deliberate commit made inside the CI image (`docker run … rust:latest … cargo x perf --pin`).
- One workspace member `smith-bench` owns every bench; it depends on any workspace crate; nothing depends on it; it is a default member so `cargo x check` type-checks benches (`cargo test --benches` is not run).
- `gungraun-runner` is a build tool installed by xtask at a pinned version matching the crate.
- Wall-clock spec targets are informational rows in the report bundle, never a gate.

## Order

1. `smith-bench` crate: library benches for frame encode and decode of 1,000 entries, `Session::from_frames`, `fork_at`, agent request build with masking, one SSE fixture per adapter, `canonical_json`, `handle_line`; binary bench for `smith --help` and `smith eval --mock`; criterion bench moved from `smith`.
2. `budgets.toml`: one `ir` pin per bench id, written by `cargo x perf --pin`.
3. `cargo x perf`: install runner, run gungraun benches with `--save-summary=json`, compare every `Ir` to its pin within ±2 %, write `target/report/perf.md`; `--pin` rewrites budgets; `cargo x bench` runs criterion with passthrough.
4. CI: `.ci/perf.kdl` per push (valgrind installed in the job); `report` bundle gains the perf section and a hyperfine row for `smith --help` and `smith eval --mock`.
5. Research note and TO_DISCUSS cleanup; deslop and compress passes.

## Interfaces

- `xtask`: `perf [--pin]`, `bench [args]`; `bootstrap` installs `gungraun-runner`.
- `smith-bench/budgets.toml`: `[<file>::<group>::<function>::<id>] ir = N`.
- `.ci/perf.kdl`.

## Risks

- Callgrind runs are slow; the set stays small and every bench runs one shot.
- Toolchain bumps shift counts; the re-pin commit names the toolchain.
- Binary benches include process start-up; that is the point for `--help`, noise for anything else.

## Exit

- `cargo x perf` green with pins; CI `perf` job observed; report bundle carries the perf section.