--- id: SMH-RESEARCH-UPU3A0HO type: research title: "Benchmarking: First Instruction Pins and Measurement Notes" --- # Benchmarking: First Instruction Pins and Measurement Notes Date: 2026-09-26. Status: evidence from implementing SMH-PLAN-2_O9OFR0; candidates, not mandates. ## Setup - Machine: AMD Ryzen 5 5600X, CachyOS, kernel 7.2.7, glibc 2.44, valgrind 3.25.1. - Toolchain: `rustc 1.98.1 (48a229cea 2026-09-01)`, bench profile (release, thin LTO, one codegen unit, line tables, not stripped). - `gungraun` 0.19.4 and `gungraun-runner` 0.19.4 (installed by `cargo x bootstrap` or `cargo x perf` through `cargo-binstall`). - Callgrind runs with `--cache-sim=no`: the gate reads only `Ir`, and cache simulation only adds run time. - `target-cpu=native` is not applied in this repo (`.cargo/config.toml` sets no `target-cpu`), so codegen is the generic x86-64 baseline. ## Pinned instruction counts Pinned by `cargo x perf --pin`, stored in `smith-bench/budgets.toml`. | bench | Ir | spread seen run to run | |---|---:|---| | `instructions::agent::run_turn::history_1000_masked` | 11270600 | ≈ 0.1 % | | `instructions::providers::sse::anthropic` | 109026 | ≤ 0.18 % | | `instructions::providers::sse::gemini` | 47234 | ≤ 0.01 % | | `instructions::providers::sse::openai` | 81236 | identical | | `instructions::rpc::handle_line::ping` | 6885 | ≈ 1.5 % (6861 to 6965) | | `instructions::session::canonical_json::nested` | 1138903 | ≤ 0.05 % | | `instructions::session::decode::linear_1000` | 11056510 | identical | | `instructions::session::decode::tool_calls_1000` | 10284262 | identical | | `instructions::session::encode::linear_1000` | 5736448 | identical | | `instructions::session::encode::tool_calls_1000` | 5133175 | identical | | `instructions::session::fork_at::linear_1000` | 232844 | ≤ 0.03 % | | `instructions::session::from_frames::linear_1000` | 1392615 | ≤ 0.07 % | | `startup::cli::smith::eval_mock` | 1156689 | ≈ 0.1 % | | `startup::cli::smith::help` | 741164 | identical | Spreads come from about fifteen runs on the machine above, direct `cargo bench` and `cargo x perf` combined. Every bench stays well inside the ±2 % band except `handle_line::ping`, whose 1.5 % spread leaves little margin. ## Why some library counts are not identical Instruction counts are exact only for code whose path does not depend on ambient randomness or time. Three sources were found in the measured code, not in the harness: - `std::collections::HashMap` and `IndexMap` (serde_json `preserve_order`) seed `RandomState` per process; probe sequences, and so instruction counts, change with the seed (`Session::from_frames`, every `serde_json` object parse, `handle_line`). - `Uuid::now_v7` takes a different branch when two IDs fall into the same millisecond (`fork_at`, `EntryFrame::new`, `run_turn`, decoders that mint a `MessageId`). - glibc `malloc` and `free` paths depend on arena state. `ping` shows it most because it is small (≈ 7,000 instructions) and dominated by one JSON parse, one `IndexMap` insert per key, and allocator calls. ## Environment findings - gungraun clears the environment of every bench but keeps `LD_LIBRARY_PATH`. `cargo run` (so every `cargo x` command) prepends library directories to it, and the dynamic loader walks each entry at start-up. The same `smith --help` binary measured 753,965 `Ir` through plain `cargo bench` and 760,523 through `cargo x perf`. `startup.rs` now sets `LD_LIBRARY_PATH` empty for the measured command; both launch paths then read 741,164. - `cargo x` children inherit `CARGO_MANIFEST_DIR=xtask` from `cargo run`. `ring`'s build script tracks that variable, so alternating `cargo x perf` with a plain `cargo build --release` rebuilds `ring`, `rustls`, `ureq`, `smith-harness`, and `smith-cli` each time (≈ 25 s here). The binary stays byte-identical; only time is lost. This predates the change and affects every `cargo x` subcommand that runs cargo. - Instruction counts depend on the toolchain, the codegen flags, and the C library. Start-up benches count the dynamic loader and libc initialization; every bench counts glibc `malloc`, `free`, `memcpy`, and `memchr`, whose implementation glibc selects at run time from the CPU features valgrind reports. Pins taken on this machine are therefore not guaranteed to hold in the `docker.io/library/rust:latest` CI image (Debian glibc). ## Run time - `cargo x perf`, warm (both builds cached): 3.2 s wall, of which the fourteen callgrind runs take about 3 s. - `cargo x perf`, cold on this machine: the release `smith` build took 2 min 42 s and the bench build 48 s, then the same 3 s of runs. - `cargo x report` with hyperfine present: 52 s here (coverage tiers dominate). ## Summary field path `--output-format=json` prints one line per bench (summary schema version `"6"`); `--save-summary=json` writes the same object to `target/gungraun/smith-bench///./summary.json`. - Bench key: top-level `module_path` (`::::`) plus `::` plus top-level `id`. - Total `Ir`: `profiles[0].summaries.total.summary.Callgrind.Ir.metrics`, which is `{"Left":{"Int":n}}` on a first run and `{"Both":[{"Int":n},{"Int":previous}]}` once gungraun has a previous run; the new value is the first `Int`. `xtask` has no JSON dependency, so `cargo x perf` reads the stdout lines with a small anchor-based extractor (`bench_ir`): the first `"module_path":"…"` and `"id":"…"`, then the first `{"Int":…}` after `"total":` → `"Ir":` → `"metrics":`. Bench names carry no escapes, and JSON string escaping keeps these anchors from matching inside the `details` strings. A line starting with `{` that the extractor cannot read fails the gate, so schema drift surfaces as a failure rather than a silent pass. Reading the per-bench `summary.json` files instead would need the same extractor plus a directory walk; the stdout lines were simpler. ## Deviations from the plan and assignment - Two gungraun bench files instead of one: `main!` accepts either library or binary groups, never both, so `instructions.rs` holds the library groups and `startup.rs` the binary group. - `cargo x perf` selects `--bench instructions --bench startup` instead of `--benches`: `--benches` would also start the criterion harness with gungraun's arguments, and cargo passes the same arguments to every selected target. The fixture library sets `bench = false` for the same reason. - Agent request build is measured through `Agent::run_turn`: `build_request` is private, and `run_turn` is its smallest public seam. The bench runs one turn over 1,000 recorded messages with a registered secret; it also records the user message and a one-delta reply, which is small next to decoding and masking 1,000 messages. The history holds no plaintext secret, which matches real sessions: recording masks input first, so masking at request build scans every message and replaces nothing. - `canonical_json` is private; the bench records a tool call with a four-level, reverse-ordered nested input through `EntryFrame::new`, which canonicalizes before encoding. - The SSE fixtures are literal copies of the adapters' conformance fixture lines: those live in `#[cfg(test)]` modules and are not reachable from another crate. - Frame encode and decode are two benches (`encode`, `decode`), each for a linear chat and a tool-call session of 1,000 entries. - The criterion bench moved from `smith/benches/message_creation.rs` to `smith-bench/benches/wall_clock.rs`. - `budgets.toml` uses one `"" = { ir = }` line per bench under `[budgets]`; no per-bench `band` key was needed, because no bench, binary benches included, varied by more than 2 % run to run. - The report's wall-clock section also exports `hyperfine.md` (hyperfine's own markdown table) next to `hyperfine.json`, so xtask needs no JSON parser for it; a `--prepare` step empties the profile before each run, because `smith eval` appends to its session file otherwise. ## Candidates - Pin in the CI image, not on a workstation, if the first CI `perf` run disagrees with these pins; then run `cargo x perf` locally in that image to compare. - Widen `handle_line::ping` only if it ever leaves the band, or measure a batch of lines so hashing noise averages out. - Remove `CARGO_MANIFEST_DIR` from the environment of cargo children in xtask to stop the `ring` rebuild ping-pong. ## Sources - [gungraun guide](https://gungraun.github.io/gungraun/latest/html/): default environment clearing, entry points, binary benchmarks, command-line arguments, machine-readable output. - `gungraun-runner` 0.19.4 source, `src/runner/args.rs`: environment clearing preserves `LD_PRELOAD` and `LD_LIBRARY_PATH`. ## Platform dependence, measured The same tree on the same rustc under Debian glibc 2.41 (the `rust:latest` image) versus CachyOS glibc 2.44: SSE decoders +9 to +11 %, `handle_line::ping` +23 %, `smith --help` −4.6 %, `eval --mock` +3.6 %, frame encode +3.3 %; the rest within 1.5 %. Startup counts include the dynamic loader; every bench includes libc `malloc` and `memcpy`, whose implementations differ per glibc release. Consequence: pins carry the platform (`# Pinned with | glibc `); the gate applies only on that platform, other hosts only report; pins are taken inside the CI image. Inside the container `setarch` cannot disable ASLR, so the runner always gets `--allow-aslr=true`; with cache simulation off the counts stayed within band across repeated container runs. Luci job containers are unprivileged: no `apt-get`. The repo therefore links with the toolchain default (mold stays a machine preference in `~/.cargo/config.toml`, whose target table composes with the repo's universal table), and the callgrind jobs build valgrind 3.25.1 from source into `target/tools/valgrind` (≈ 3 min, Luci-cached on the job file). Pins were re-taken with that exact setup; linker choice shifts counts, so pins never move without the platform line.