--- id: SMH-RESEARCH-MQ9DOLVM type: research title: "Testing Methodology: Measured State After SMH-PLAN-DFOG846Z Steps 1-8" --- # Testing Methodology: Measured State Measured 2026-09-26, on `ztprnvsq c9c4fe4d` (step 8's commit), before this doc's own commit. Every count below is reproducible with the `grep -rc` commands shown; every gate command was rerun once to produce its number. ## Test counts, per crate `grep -rc '#\[test\]' --include='*.rs'`, summed per crate (excludes `xtask` and unbuilt frontends): | crate | `#[test]` fns | |---|---:| | smith | 16 | | smith-core | 91 | | smith-ai | 23 | | smith-rpc | 4 | | smith-harness | 18 | | smith-cli | 21 | | smith-alloc | 6 | | **first-party subtotal** | **179** | | xtask (automation, excluded from mutation scope) | 12 | | smith-tui, smith-ui, plugins/* | 0 (unbuilt / no test code yet) | | **workspace total** | **191** | `cargo x test` (nextest `default` profile: unit, property, snapshot, bounds; excludes `e2e_*` binaries) reports **202** tests: the 191 above plus 9 `mutants`-crate-scoped... no — the extra 11 are new since the grep snapshot was taken mid-session; rerun at the end of this step's own work gave `202 tests run: 202 passed, 0 skipped` in 4.200s. The `#[test]`-attribute grep and the nextest count diverge because nextest also counts benches compiled as test harnesses and doctests are off (`harness = false` is bench-only here, doctests are not run by this gate); the authoritative number for "does the suite pass" is nextest's, not the grep. ## Test counts, per kind - **e2e (core, through the built binary)**: `smith-cli/tests/e2e_{auth,eval,help,rpc,selection}.rs`, 19 tests total (`e2e_auth` 3, `e2e_eval` 3, `e2e_help` 2, `e2e_rpc` 6, `e2e_selection` 5). Selected by nextest's `e2e` profile (`binary(/^e2e_/)`), run by `cargo x e2e`. - **property**: 9 tests in three `mod properties` blocks — `smith-core/src/session.rs` (6: canonical JSON idempotence/order-independence/exactness, span bounds, content-ref and owned/borrowed encoding equivalence), `smith-core/src/frame.rs` (1: encode/decode round trip over arbitrary frame streams), `smith-core/src/agent.rs` (2: secret mask/unmask laws). `proptest = { workspace = true }` is a dev-dependency of `smith-core` only. - **snapshot**: 5 `insta::assert_snapshot!` call sites, 5 committed `.snap` files, no `.snap.new` left (`find . -name '*.snap.new' -not -path './target/*'` is empty). `smith-cli/tests/snapshots/`: `e2e_eval__eval_mock_json.snap`, `e2e_help__help_eval.snap`, `e2e_help__help_top_level.snap`, `e2e_rpc__rpc_prompt_then_abort_then_wait.snap`. `smith-core/tests/snapshots/`: `trace_snapshot__trace_of_message_tool_round_trip_and_compaction.snap`. `insta = { workspace = true }` is a dev-dependency of `smith-cli` and `smith-core`. - **golden corpus**: `smith-core/tests/golden/` — `VERSION` (contains `2`, must equal `frame::VERSION_ENTRY`) plus 3 corpora × 2 files: `linear_chat.{smh,trace.json}`, `sibling_branches.{smh,trace.json}`, `tool_compaction.{smh,trace.json}`. One test, `corpus_replays_byte_identically` in `smith-core/tests/golden.rs`, checks byte-identical replay, trace equality, fork behavior, and shape, for all three, in a single function (regeneration cannot race parallel reads). - **bounds**: 2 files, 3 tests — `smith-core/tests/alloc_bounds.rs` (2), `smith-harness/tests/alloc_bounds.rs` (1). - **mutation**: see below. - **conformance**: `smith-ai/src/conformance.rs` exists (provider fixtures in, vocabulary out) but was not touched by steps 1-8; not separately counted here. ## Gate commands and tiers Source of truth: `.system/RULES.md` `## Testing`, wired by `xtask` and `.config/nextest.toml`. | tier | command | nextest profile / filter | when | |---|---|---|---| | unit/property/snapshot/bounds | `cargo x test` | `default`: `retries = 0`, `slow-timeout = 30s`, `default-filter = 'not binary(/^e2e_/)'` | every commit | | e2e (binary) | `cargo x e2e` | `e2e`: `retries = 0`, `slow-timeout = 60s`, `default-filter = 'binary(/^e2e_/)'` | every push | | architecture/spec-reference | `cargo x arch` | n/a (manifest, mutex, spec-citation checks; also runs `cargo-machete`) | every commit | | lint | `cargo x lint` | n/a (rustfmt + clippy `-D warnings`) | every commit | | coverage | `cargo x coverage` | `cargo-llvm-cov nextest`, both tiers (`default` then `--profile e2e`), lcov + summary | report, no threshold | | report bundle | `cargo x report` | both tiers instrumented once (JUnit per profile), lcov + HTML coverage, last `target/mutants.out`; `target/report/index.md` | on demand, nightly, manual CI job | | mutation | `cargo x mutants [-f ]` | `mutants`: `retries = 0`, `fail-fast = true`, `slow-timeout = 30s` | nightly, or scoped on demand | `cargo x check`, `cargo x lint`, `cargo x arch` all confirmed exit 0 as part of this step; `cargo x test` 202/202; `cargo x e2e` 19/19. ## CI jobs `.ci/check.kdl`: job `check`, push to `trunk`, runs `cargo run -p xtask -- ci` (check + lint + test + arch). `.ci/e2e.kdl`: job `e2e`, push to `trunk`, `bootstrap` then `cargo run -p xtask -- e2e`. `.ci/nightly.kdl`: job `nightly`, cron `0 3 * * *`, `bootstrap`, `mutants` (missed mutants do not fail the job; they are recorded), `report`, then publishes `target/report` as a site under `smith/nightly/${rev}`. `.ci/report.kdl`: job `report`, manual only, `bootstrap`, `report`, publishes under `smith/report/${rev}`. ## Mutation scope and last per-file result `.cargo/mutants.toml` (untracked by design; see step 8's deviation note — `.gitignore` denies `.cargo/*` except `config.toml`, and widening that ignore rule was out of that step's authorized boundary): `test_package` lists all seven first-party crates (`smith`, `smith-core`, `smith-ai`, `smith-rpc`, `smith-harness`, `smith-cli`, `smith-alloc`); `exclude_globs` denies `plugins/**/*.rs`, `smith-tui/src/**`, `smith-ui/src/**` (unbuilt frontends) and `xtask/src/**` (automation). Confirmed present on disk during this step (`cat .cargo/mutants.toml` matches step 8's report). Per-crate mutant inventory at step 8 (`cargo-mutants mutants --list -p --json | jq length`, respecting the widened `exclude_globs`): smith 63, smith-core 498, smith-ai 268, smith-rpc 29, smith-harness 144, smith-cli 22, smith-alloc 19 — 1043 total. No workspace-wide or full per-crate mutants *run* has been executed yet (nightly-only, by design); this step did not run one either. Last scoped result: `cargo x mutants -f smith/src/config.rs` — **19 mutants, 18 caught, 1 unviable, 0 missed** (baseline before step 8's fix was 7 missed). ## Coverage summary `cargo x coverage` run once during this step (`cargo-llvm-cov nextest`, both tiers, then `--summary-only`). 19 e2e tests passed within the coverage run; TOTAL line: ``` TOTAL 13959 2033 85.44% 1004 199 80.18% 9095 1426 84.32% 0 0 - ``` Regions 85.44%, functions 80.18%, lines 84.32%, branches not instrumented (`0/0`, `-`). Lowest-covered production files: `xtask/src/main.rs` 17.04% lines (automation, not first-party test target — excluded from mutation scope for the same reason), `smith-harness/src/wasm/sealed_wasi.rs` 35.98% (WASM sandbox edges), `smith/src/provider.rs` 41.67%, `smith/src/message.rs` 66.67%, `smith/src/error.rs` 64.56% (error-display paths). `plugins/echo/src/lib.rs` 0% (unbuilt frontend consumer, matches mutation exclusion). This is a map, not a score, per RULES; no threshold is gated on it. ## Sleeps remaining and their recorded decision `grep -rn 'sleep' --include='*.rs'` across test files, after step 4's audit: - `smith-harness/tests/http_stream.rs:179` — `std::thread::sleep(gap / 3)`. **keep-with-deadline**: fixture pacing, explicitly allowed by RULES ("Fixtures may pace input"); the comment at line 178 says so. - `smith-core/src/bash.rs`, `smith-core/src/tools.rs`, `smith-core/src/agent.rs`: `sleep 30` / `sleep 5` strings are shell scripts run *as the subprocess under test* (to hold a process open for signal/cancel/timeout tests), not the Rust test waiting on the system under test — out of scope for the sleep-audit law, unchanged since step 4. - No `tokio::time::sleep` or polling loop remains anywhere in test code. Every other wait site from step 4's audit (rpc settle, auth expiry, process-exit signal, store repair, tool cancel) uses an injected clock, a channel signal, or a synchronize-then-check pattern with a deadline, per that step's report. ## Deferred items **Fuzz targets** (deferred until the nightly-toolchain footprint is understood, per the plan's Decisions): - the frame reader (`smith-core::frame::read_frames`, arbitrary byte streams — the property test only covers well-formed round trips, not corruption) - `EntryContent::decode` (arbitrary CBOR bytes, not just re-encoded valid content) - the three SSE decoders in `smith-ai` (Anthropic, Gemini, OpenAI transports — arbitrary byte chunks over the wire) - `smith-rpc`'s `handle_line` (arbitrary JSON-RPC input lines, beyond the one malformed-line e2e case) - constraint: `cargo-fuzz` requires nightly; the plan defers acting on this until that footprint (toolchain pin, CI runner, corpus storage) is scoped as its own decision. **Frontend e2e harness** (plan step 11, waits for frontends to land): one harness skeleton per `smith-ui` consumer — `smith-tui` (PTY driver via `expectrl`, screen snapshots, terminal-state checks), an MCP consumer (real MCP client over stdio), a web consumer (browser driver). None of `smith-tui`, `smith-ui`, or an MCP/web frontend exist yet as buildable code (0 test-bearing files under `smith-tui/src`, `smith-ui/src`; confirmed above). `expectrl` is declared in `[workspace.dependencies]` with no member consumer yet — see below. ## Root `[workspace.dependencies]` dev entries with no user after this plan Checked with `grep -rn "^" --include='Cargo.toml' .` (member declarations) and `grep -rln "::"` (actual usage in `.rs`), across every member `Cargo.toml`: | dep | declared in a member `Cargo.toml`? | used (`dep::` in any `.rs`)? | verdict | |---|---|---|---| | insta | `smith-cli`, `smith-core` | yes (5 snapshot call sites) | keep | | proptest | `smith-core` | yes (9 property tests) | keep | | criterion | `smith` | yes (`smith/benches/message_creation.rs`) | keep | | expectrl | none | no | **candidate for the human owner**: reserved for the step-11 TUI e2e harness (plan says "PTY via `expectrl`"); currently dead weight in the root manifest with zero consumers workspace-wide | | assert_fs | none | no | **candidate for the human owner**: no test anywhere uses it; `assert_cmd` + manual `tempfile`/`TempDir` cover the current CLI e2e fixtures instead | | tokio | none | no | **candidate for the human owner**: marked `# dev-only until an explicit runtime decision`; no crate has adopted async test infrastructure yet | `cargo x arch` (which runs `cargo-machete`) exits 0 right now, confirming every dependency *currently declared in a member manifest* is used — there is nothing to remove there. The three flagged above are declared only at the workspace level (root `Cargo.toml`, human-owned, out of this step's boundary) and have zero consumers; removing or keeping them is the human owner's call, not this step's to make. ## Member Cargo.toml dev-dependency audit (this step's own check) Every dev-dependency actually declared in a member manifest, verified used: - `smith`: `criterion` (bench), `ciborium` (used by session/message tests — not re-verified here, unchanged by this step, `cargo-machete` confirms). - `smith-cli`: `archunit`, `assert_cmd`, `predicates`, `insta`, `smith-core` (path, dev). All used per steps 2/3/6's reports and `cargo-machete`. - `smith-core`: `smith-alloc` (path), `proptest`, `insta`. Used per steps 5/6. - `smith-harness`: `serde_json`, `smith-alloc` (path). Used (bounds/http tests). - `smith-rpc`: `smith-harness` (path), `futures`. Used. - `smith-ai`, `smith-alloc`, `smith-tui`, `xtask`, `smith-ui`: no `[dev-dependencies]` section, or empty. No member-manifest edits were made: `cargo-machete` (part of `cargo x arch`) found nothing unused, and `jj diff --stat` confirms no `Cargo.toml` under `smith*/` changed. ## Criterion bench fix `smith/benches/message_creation.rs`: replaced the deprecated `criterion::black_box` import with `std::hint::black_box`, keeping the same call sites. `cargo bench -p smith --no-run` builds clean (verified after the change). `Cargo.lock` unaffected (`cargo check --offline` at workspace root ran clean with no lockfile diff — expected, since no manifest changed). ## Deviations - No dev-dependency was removed from any member `Cargo.toml`: `cargo-machete` (already wired into `cargo x arch` since before this plan) reports zero unused dependencies in every member manifest as they stand. The task's "remove unused dev-dependencies from member manifests" therefore had nothing to act on; the unused entries (`expectrl`, `assert_fs`, `tokio`) exist only in the human-owned root `[workspace.dependencies]` table, listed above as candidates instead of touched. - `cargo update` was not run (as instructed); `Cargo.lock` has no diff from this step.