Luigit
repositories / smith

smith

There are many coding harnesses - but this one is fast

owned by admin

.system/research/SMH-RESEARCH-UPU3A0HO-benchmarking/index.md

Raw
Rendered preview

id: SMH-RESEARCH-UPU3A0HO type: research title: "Benchmarking: First Instruction Pins and Measurement Notes"

Benchmarking: First Instruction Pins and Measurement Notes

Date: 2026-09-26. Status: evidence from implementing SMH-PLAN-2_O9OFR0; candidates, not mandates.

Setup

  • Machine: AMD Ryzen 5 5600X, CachyOS, kernel 7.2.7, glibc 2.44, valgrind 3.25.1.
  • Toolchain: rustc 1.98.1 (48a229cea 2026-09-01), bench profile (release, thin LTO, one codegen unit, line tables, not stripped).
  • gungraun 0.19.4 and gungraun-runner 0.19.4 (installed by cargo x bootstrap or cargo x perf through cargo-binstall).
  • Callgrind runs with --cache-sim=no: the gate reads only Ir, and cache simulation only adds run time.
  • target-cpu=native is not applied in this repo (.cargo/config.toml sets no target-cpu), so codegen is the generic x86-64 baseline.

Pinned instruction counts

Pinned by cargo x perf --pin, stored in smith-bench/budgets.toml.

bench Ir spread seen run to run
instructions::agent::run_turn::history_1000_masked 11270600 ≈ 0.1 %
instructions::providers::sse::anthropic 109026 ≤ 0.18 %
instructions::providers::sse::gemini 47234 ≤ 0.01 %
instructions::providers::sse::openai 81236 identical
instructions::rpc::handle_line::ping 6885 ≈ 1.5 % (6861 to 6965)
instructions::session::canonical_json::nested 1138903 ≤ 0.05 %
instructions::session::decode::linear_1000 11056510 identical
instructions::session::decode::tool_calls_1000 10284262 identical
instructions::session::encode::linear_1000 5736448 identical
instructions::session::encode::tool_calls_1000 5133175 identical
instructions::session::fork_at::linear_1000 232844 ≤ 0.03 %
instructions::session::from_frames::linear_1000 1392615 ≤ 0.07 %
startup::cli::smith::eval_mock 1156689 ≈ 0.1 %
startup::cli::smith::help 741164 identical

Spreads come from about fifteen runs on the machine above, direct cargo bench and cargo x perf combined. Every bench stays well inside the ±2 % band except handle_line::ping, whose 1.5 % spread leaves little margin.

Why some library counts are not identical

Instruction counts are exact only for code whose path does not depend on ambient randomness or time. Three sources were found in the measured code, not in the harness:

  • std::collections::HashMap and IndexMap (serde_json preserve_order) seed RandomState per process; probe sequences, and so instruction counts, change with the seed (Session::from_frames, every serde_json object parse, handle_line).
  • Uuid::now_v7 takes a different branch when two IDs fall into the same millisecond (fork_at, EntryFrame::new, run_turn, decoders that mint a MessageId).
  • glibc malloc and free paths depend on arena state.

ping shows it most because it is small (≈ 7,000 instructions) and dominated by one JSON parse, one IndexMap insert per key, and allocator calls.

Environment findings

  • gungraun clears the environment of every bench but keeps LD_LIBRARY_PATH. cargo run (so every cargo x command) prepends library directories to it, and the dynamic loader walks each entry at start-up. The same smith --help binary measured 753,965 Ir through plain cargo bench and 760,523 through cargo x perf. startup.rs now sets LD_LIBRARY_PATH empty for the measured command; both launch paths then read 741,164.
  • cargo x children inherit CARGO_MANIFEST_DIR=xtask from cargo run. ring's build script tracks that variable, so alternating cargo x perf with a plain cargo build --release rebuilds ring, rustls, ureq, smith-harness, and smith-cli each time (≈ 25 s here). The binary stays byte-identical; only time is lost. This predates the change and affects every cargo x subcommand that runs cargo.
  • Instruction counts depend on the toolchain, the codegen flags, and the C library. Start-up benches count the dynamic loader and libc initialization; every bench counts glibc malloc, free, memcpy, and memchr, whose implementation glibc selects at run time from the CPU features valgrind reports. Pins taken on this machine are therefore not guaranteed to hold in the docker.io/library/rust:latest CI image (Debian glibc).

Run time

  • cargo x perf, warm (both builds cached): 3.2 s wall, of which the fourteen callgrind runs take about 3 s.
  • cargo x perf, cold on this machine: the release smith build took 2 min 42 s and the bench build 48 s, then the same 3 s of runs.
  • cargo x report with hyperfine present: 52 s here (coverage tiers dominate).

Summary field path

--output-format=json prints one line per bench (summary schema version "6"); --save-summary=json writes the same object to target/gungraun/smith-bench/<file>/<group>/<function>.<id>/summary.json.

  • Bench key: top-level module_path (<file>::<group>::<function>) plus :: plus top-level id.
  • Total Ir: profiles[0].summaries.total.summary.Callgrind.Ir.metrics, which is {"Left":{"Int":n}} on a first run and {"Both":[{"Int":n},{"Int":previous}]} once gungraun has a previous run; the new value is the first Int.

xtask has no JSON dependency, so cargo x perf reads the stdout lines with a small anchor-based extractor (bench_ir): the first "module_path":"…" and "id":"…", then the first {"Int":…} after "total": → "Ir": → "metrics":. Bench names carry no escapes, and JSON string escaping keeps these anchors from matching inside the details strings. A line starting with { that the extractor cannot read fails the gate, so schema drift surfaces as a failure rather than a silent pass. Reading the per-bench summary.json files instead would need the same extractor plus a directory walk; the stdout lines were simpler.

Deviations from the plan and assignment

  • Two gungraun bench files instead of one: main! accepts either library or binary groups, never both, so instructions.rs holds the library groups and startup.rs the binary group.
  • cargo x perf selects --bench instructions --bench startup instead of --benches: --benches would also start the criterion harness with gungraun's arguments, and cargo passes the same arguments to every selected target. The fixture library sets bench = false for the same reason.
  • Agent request build is measured through Agent::run_turn: build_request is private, and run_turn is its smallest public seam. The bench runs one turn over 1,000 recorded messages with a registered secret; it also records the user message and a one-delta reply, which is small next to decoding and masking 1,000 messages. The history holds no plaintext secret, which matches real sessions: recording masks input first, so masking at request build scans every message and replaces nothing.
  • canonical_json is private; the bench records a tool call with a four-level, reverse-ordered nested input through EntryFrame::new, which canonicalizes before encoding.
  • The SSE fixtures are literal copies of the adapters' conformance fixture lines: those live in #[cfg(test)] modules and are not reachable from another crate.
  • Frame encode and decode are two benches (encode, decode), each for a linear chat and a tool-call session of 1,000 entries.
  • The criterion bench moved from smith/benches/message_creation.rs to smith-bench/benches/wall_clock.rs.
  • budgets.toml uses one "<bench>" = { ir = <n> } line per bench under [budgets]; no per-bench band key was needed, because no bench, binary benches included, varied by more than 2 % run to run.
  • The report's wall-clock section also exports hyperfine.md (hyperfine's own markdown table) next to hyperfine.json, so xtask needs no JSON parser for it; a --prepare step empties the profile before each run, because smith eval appends to its session file otherwise.

Candidates

  • Pin in the CI image, not on a workstation, if the first CI perf run disagrees with these pins; then run cargo x perf locally in that image to compare.
  • Widen handle_line::ping only if it ever leaves the band, or measure a batch of lines so hashing noise averages out.
  • Remove CARGO_MANIFEST_DIR from the environment of cargo children in xtask to stop the ring rebuild ping-pong.

Sources

  • gungraun guide: default environment clearing, entry points, binary benchmarks, command-line arguments, machine-readable output.
  • gungraun-runner 0.19.4 source, src/runner/args.rs: environment clearing preserves LD_PRELOAD and LD_LIBRARY_PATH.

Platform dependence, measured

The same tree on the same rustc under Debian glibc 2.41 (the rust:latest image) versus CachyOS glibc 2.44: SSE decoders +9 to +11 %, handle_line::ping +23 %, smith --help −4.6 %, eval --mock +3.6 %, frame encode +3.3 %; the rest within 1.5 %. Startup counts include the dynamic loader; every bench includes libc malloc and memcpy, whose implementations differ per glibc release. Consequence: pins carry the platform (# Pinned with <rustc> | glibc <v>); the gate applies only on that platform, other hosts only report; pins are taken inside the CI image. Inside the container setarch cannot disable ASLR, so the runner always gets --allow-aslr=true; with cache simulation off the counts stayed within band across repeated container runs. Luci job containers are unprivileged: no apt-get. The repo therefore links with the toolchain default (mold stays a machine preference in ~/.cargo/config.toml, whose target table composes with the repo's universal table), and the callgrind jobs build valgrind 3.25.1 from source into target/tools/valgrind (≈ 3 min, Luci-cached on the job file). Pins were re-taken with that exact setup; linker choice shifts counts, so pins never move without the platform line.

---
id: SMH-RESEARCH-UPU3A0HO
type: research
title: "Benchmarking: First Instruction Pins and Measurement Notes"
---

# Benchmarking: First Instruction Pins and Measurement Notes

Date: 2026-09-26.
Status: evidence from implementing SMH-PLAN-2_O9OFR0; candidates, not mandates.

## Setup

- Machine: AMD Ryzen 5 5600X, CachyOS, kernel 7.2.7, glibc 2.44, valgrind 3.25.1.
- Toolchain: `rustc 1.98.1 (48a229cea 2026-09-01)`, bench profile (release, thin LTO, one codegen unit, line tables, not stripped).
- `gungraun` 0.19.4 and `gungraun-runner` 0.19.4 (installed by `cargo x bootstrap` or `cargo x perf` through `cargo-binstall`).
- Callgrind runs with `--cache-sim=no`: the gate reads only `Ir`, and cache simulation only adds run time.
- `target-cpu=native` is not applied in this repo (`.cargo/config.toml` sets no `target-cpu`), so codegen is the generic x86-64 baseline.

## Pinned instruction counts

Pinned by `cargo x perf --pin`, stored in `smith-bench/budgets.toml`.

| bench | Ir | spread seen run to run |
|---|---:|---|
| `instructions::agent::run_turn::history_1000_masked` | 11270600 | ≈ 0.1 % |
| `instructions::providers::sse::anthropic` | 109026 | ≤ 0.18 % |
| `instructions::providers::sse::gemini` | 47234 | ≤ 0.01 % |
| `instructions::providers::sse::openai` | 81236 | identical |
| `instructions::rpc::handle_line::ping` | 6885 | ≈ 1.5 % (6861 to 6965) |
| `instructions::session::canonical_json::nested` | 1138903 | ≤ 0.05 % |
| `instructions::session::decode::linear_1000` | 11056510 | identical |
| `instructions::session::decode::tool_calls_1000` | 10284262 | identical |
| `instructions::session::encode::linear_1000` | 5736448 | identical |
| `instructions::session::encode::tool_calls_1000` | 5133175 | identical |
| `instructions::session::fork_at::linear_1000` | 232844 | ≤ 0.03 % |
| `instructions::session::from_frames::linear_1000` | 1392615 | ≤ 0.07 % |
| `startup::cli::smith::eval_mock` | 1156689 | ≈ 0.1 % |
| `startup::cli::smith::help` | 741164 | identical |

Spreads come from about fifteen runs on the machine above, direct `cargo bench` and `cargo x perf` combined.
Every bench stays well inside the ±2 % band except `handle_line::ping`, whose 1.5 % spread leaves little margin.

## Why some library counts are not identical

Instruction counts are exact only for code whose path does not depend on ambient randomness or time.
Three sources were found in the measured code, not in the harness:

- `std::collections::HashMap` and `IndexMap` (serde_json `preserve_order`) seed `RandomState` per process; probe sequences, and so instruction counts, change with the seed (`Session::from_frames`, every `serde_json` object parse, `handle_line`).
- `Uuid::now_v7` takes a different branch when two IDs fall into the same millisecond (`fork_at`, `EntryFrame::new`, `run_turn`, decoders that mint a `MessageId`).
- glibc `malloc` and `free` paths depend on arena state.

`ping` shows it most because it is small (≈ 7,000 instructions) and dominated by one JSON parse, one `IndexMap` insert per key, and allocator calls.

## Environment findings

- gungraun clears the environment of every bench but keeps `LD_LIBRARY_PATH`.
  `cargo run` (so every `cargo x` command) prepends library directories to it, and the dynamic loader walks each entry at start-up.
  The same `smith --help` binary measured 753,965 `Ir` through plain `cargo bench` and 760,523 through `cargo x perf`.
  `startup.rs` now sets `LD_LIBRARY_PATH` empty for the measured command; both launch paths then read 741,164.
- `cargo x` children inherit `CARGO_MANIFEST_DIR=xtask` from `cargo run`.
  `ring`'s build script tracks that variable, so alternating `cargo x perf` with a plain `cargo build --release` rebuilds `ring`, `rustls`, `ureq`, `smith-harness`, and `smith-cli` each time (≈ 25 s here).
  The binary stays byte-identical; only time is lost.
  This predates the change and affects every `cargo x` subcommand that runs cargo.
- Instruction counts depend on the toolchain, the codegen flags, and the C library.
  Start-up benches count the dynamic loader and libc initialization; every bench counts glibc `malloc`, `free`, `memcpy`, and `memchr`, whose implementation glibc selects at run time from the CPU features valgrind reports.
  Pins taken on this machine are therefore not guaranteed to hold in the `docker.io/library/rust:latest` CI image (Debian glibc).

## Run time

- `cargo x perf`, warm (both builds cached): 3.2 s wall, of which the fourteen callgrind runs take about 3 s.
- `cargo x perf`, cold on this machine: the release `smith` build took 2 min 42 s and the bench build 48 s, then the same 3 s of runs.
- `cargo x report` with hyperfine present: 52 s here (coverage tiers dominate).

## Summary field path

`--output-format=json` prints one line per bench (summary schema version `"6"`); `--save-summary=json` writes the same object to `target/gungraun/smith-bench/<file>/<group>/<function>.<id>/summary.json`.

- Bench key: top-level `module_path` (`<file>::<group>::<function>`) plus `::` plus top-level `id`.
- Total `Ir`: `profiles[0].summaries.total.summary.Callgrind.Ir.metrics`, which is `{"Left":{"Int":n}}` on a first run and `{"Both":[{"Int":n},{"Int":previous}]}` once gungraun has a previous run; the new value is the first `Int`.

`xtask` has no JSON dependency, so `cargo x perf` reads the stdout lines with a small anchor-based extractor (`bench_ir`): the first `"module_path":"…"` and `"id":"…"`, then the first `{"Int":…}` after `"total":` → `"Ir":` → `"metrics":`.
Bench names carry no escapes, and JSON string escaping keeps these anchors from matching inside the `details` strings.
A line starting with `{` that the extractor cannot read fails the gate, so schema drift surfaces as a failure rather than a silent pass.
Reading the per-bench `summary.json` files instead would need the same extractor plus a directory walk; the stdout lines were simpler.

## Deviations from the plan and assignment

- Two gungraun bench files instead of one: `main!` accepts either library or binary groups, never both, so `instructions.rs` holds the library groups and `startup.rs` the binary group.
- `cargo x perf` selects `--bench instructions --bench startup` instead of `--benches`: `--benches` would also start the criterion harness with gungraun's arguments, and cargo passes the same arguments to every selected target.
  The fixture library sets `bench = false` for the same reason.
- Agent request build is measured through `Agent::run_turn`: `build_request` is private, and `run_turn` is its smallest public seam.
  The bench runs one turn over 1,000 recorded messages with a registered secret; it also records the user message and a one-delta reply, which is small next to decoding and masking 1,000 messages.
  The history holds no plaintext secret, which matches real sessions: recording masks input first, so masking at request build scans every message and replaces nothing.
- `canonical_json` is private; the bench records a tool call with a four-level, reverse-ordered nested input through `EntryFrame::new`, which canonicalizes before encoding.
- The SSE fixtures are literal copies of the adapters' conformance fixture lines: those live in `#[cfg(test)]` modules and are not reachable from another crate.
- Frame encode and decode are two benches (`encode`, `decode`), each for a linear chat and a tool-call session of 1,000 entries.
- The criterion bench moved from `smith/benches/message_creation.rs` to `smith-bench/benches/wall_clock.rs`.
- `budgets.toml` uses one `"<bench>" = { ir = <n> }` line per bench under `[budgets]`; no per-bench `band` key was needed, because no bench, binary benches included, varied by more than 2 % run to run.
- The report's wall-clock section also exports `hyperfine.md` (hyperfine's own markdown table) next to `hyperfine.json`, so xtask needs no JSON parser for it; a `--prepare` step empties the profile before each run, because `smith eval` appends to its session file otherwise.

## Candidates

- Pin in the CI image, not on a workstation, if the first CI `perf` run disagrees with these pins; then run `cargo x perf` locally in that image to compare.
- Widen `handle_line::ping` only if it ever leaves the band, or measure a batch of lines so hashing noise averages out.
- Remove `CARGO_MANIFEST_DIR` from the environment of cargo children in xtask to stop the `ring` rebuild ping-pong.

## Sources

- [gungraun guide](https://gungraun.github.io/gungraun/latest/html/): default environment clearing, entry points, binary benchmarks, command-line arguments, machine-readable output.
- `gungraun-runner` 0.19.4 source, `src/runner/args.rs`: environment clearing preserves `LD_PRELOAD` and `LD_LIBRARY_PATH`.

## Platform dependence, measured

The same tree on the same rustc under Debian glibc 2.41 (the `rust:latest` image) versus CachyOS glibc 2.44: SSE decoders +9 to +11 %, `handle_line::ping` +23 %, `smith --help` −4.6 %, `eval --mock` +3.6 %, frame encode +3.3 %; the rest within 1.5 %.
Startup counts include the dynamic loader; every bench includes libc `malloc` and `memcpy`, whose implementations differ per glibc release.
Consequence: pins carry the platform (`# Pinned with <rustc> | glibc <v>`); the gate applies only on that platform, other hosts only report; pins are taken inside the CI image.
Inside the container `setarch` cannot disable ASLR, so the runner always gets `--allow-aslr=true`; with cache simulation off the counts stayed within band across repeated container runs.
Luci job containers are unprivileged: no `apt-get`. The repo therefore links with the toolchain default (mold stays a machine preference in `~/.cargo/config.toml`, whose target table composes with the repo's universal table), and the callgrind jobs build valgrind 3.25.1 from source into `target/tools/valgrind` (≈ 3 min, Luci-cached on the job file). Pins were re-taken with that exact setup; linker choice shifts counts, so pins never move without the platform line.