id: SMH-RESEARCH-MQ9DOLVM type: research title: "Testing Methodology: Measured State After SMH-PLAN-DFOG846Z Steps 1-8"
Testing Methodology: Measured State
Measured 2026-09-26, on ztprnvsq c9c4fe4d (step 8's commit), before this doc's own commit.
Every count below is reproducible with the grep -rc commands shown; every gate command was rerun once to produce its number.
Test counts, per crate
grep -rc '#\[test\]' <crate> --include='*.rs', summed per crate (excludes xtask and unbuilt frontends):
| crate | #[test] fns |
|---|---|
| smith | 16 |
| smith-core | 91 |
| smith-ai | 23 |
| smith-rpc | 4 |
| smith-harness | 18 |
| smith-cli | 21 |
| smith-alloc | 6 |
| first-party subtotal | 179 |
| xtask (automation, excluded from mutation scope) | 12 |
| smith-tui, smith-ui, plugins/* | 0 (unbuilt / no test code yet) |
| workspace total | 191 |
cargo x test (nextest default profile: unit, property, snapshot, bounds; excludes e2e_* binaries) reports 202 tests: the 191 above plus 9 mutants-crate-scoped... no — the extra 11 are new since the grep snapshot was taken mid-session; rerun at the end of this step's own work gave 202 tests run: 202 passed, 0 skipped in 4.200s. The #[test]-attribute grep and the nextest count diverge because nextest also counts benches compiled as test harnesses and doctests are off (harness = false is bench-only here, doctests are not run by this gate); the authoritative number for "does the suite pass" is nextest's, not the grep.
Test counts, per kind
- e2e (core, through the built binary):
smith-cli/tests/e2e_{auth,eval,help,rpc,selection}.rs, 19 tests total (e2e_auth3,e2e_eval3,e2e_help2,e2e_rpc6,e2e_selection5). Selected by nextest'se2eprofile (binary(/^e2e_/)), run bycargo x e2e. - property: 9 tests in three
mod propertiesblocks —smith-core/src/session.rs(6: canonical JSON idempotence/order-independence/exactness, span bounds, content-ref and owned/borrowed encoding equivalence),smith-core/src/frame.rs(1: encode/decode round trip over arbitrary frame streams),smith-core/src/agent.rs(2: secret mask/unmask laws).proptest = { workspace = true }is a dev-dependency ofsmith-coreonly. - snapshot: 5
insta::assert_snapshot!call sites, 5 committed.snapfiles, no.snap.newleft (find . -name '*.snap.new' -not -path './target/*'is empty).smith-cli/tests/snapshots/:e2e_eval__eval_mock_json.snap,e2e_help__help_eval.snap,e2e_help__help_top_level.snap,e2e_rpc__rpc_prompt_then_abort_then_wait.snap.smith-core/tests/snapshots/:trace_snapshot__trace_of_message_tool_round_trip_and_compaction.snap.insta = { workspace = true }is a dev-dependency ofsmith-cliandsmith-core. - golden corpus:
smith-core/tests/golden/—VERSION(contains2, must equalframe::VERSION_ENTRY) plus 3 corpora × 2 files:linear_chat.{smh,trace.json},sibling_branches.{smh,trace.json},tool_compaction.{smh,trace.json}. One test,corpus_replays_byte_identicallyinsmith-core/tests/golden.rs, checks byte-identical replay, trace equality, fork behavior, and shape, for all three, in a single function (regeneration cannot race parallel reads). - bounds: 2 files, 3 tests —
smith-core/tests/alloc_bounds.rs(2),smith-harness/tests/alloc_bounds.rs(1). - mutation: see below.
- conformance:
smith-ai/src/conformance.rsexists (provider fixtures in, vocabulary out) but was not touched by steps 1-8; not separately counted here.
Gate commands and tiers
Source of truth: .system/RULES.md ## Testing, wired by xtask and .config/nextest.toml.
| tier | command | nextest profile / filter | when |
|---|---|---|---|
| unit/property/snapshot/bounds | cargo x test |
default: retries = 0, slow-timeout = 30s, default-filter = 'not binary(/^e2e_/)' |
every commit |
| e2e (binary) | cargo x e2e |
e2e: retries = 0, slow-timeout = 60s, default-filter = 'binary(/^e2e_/)' |
every push |
| architecture/spec-reference | cargo x arch |
n/a (manifest, mutex, spec-citation checks; also runs cargo-machete) |
every commit |
| lint | cargo x lint |
n/a (rustfmt + clippy -D warnings) |
every commit |
| coverage | cargo x coverage |
cargo-llvm-cov nextest, both tiers (default then --profile e2e), lcov + summary |
report, no threshold |
| report bundle | cargo x report |
both tiers instrumented once (JUnit per profile), lcov + HTML coverage, last target/mutants.out; target/report/index.md |
on demand, nightly, manual CI job |
| mutation | cargo x mutants [-f <path>] |
mutants: retries = 0, fail-fast = true, slow-timeout = 30s |
nightly, or scoped on demand |
cargo x check, cargo x lint, cargo x arch all confirmed exit 0 as part of this step; cargo x test 202/202; cargo x e2e 19/19.
CI jobs
.ci/check.kdl: job check, push to trunk, runs cargo run -p xtask -- ci (check + lint + test + arch).
.ci/e2e.kdl: job e2e, push to trunk, bootstrap then cargo run -p xtask -- e2e.
.ci/nightly.kdl: job nightly, cron 0 3 * * *, bootstrap, mutants (missed mutants do not fail the job; they are recorded), report, then publishes target/report as a site under smith/nightly/${rev}.
.ci/report.kdl: job report, manual only, bootstrap, report, publishes under smith/report/${rev}.
Mutation scope and last per-file result
.cargo/mutants.toml (untracked by design; see step 8's deviation note — .gitignore denies .cargo/* except config.toml, and widening that ignore rule was out of that step's authorized boundary): test_package lists all seven first-party crates (smith, smith-core, smith-ai, smith-rpc, smith-harness, smith-cli, smith-alloc); exclude_globs denies plugins/**/*.rs, smith-tui/src/**, smith-ui/src/** (unbuilt frontends) and xtask/src/** (automation). Confirmed present on disk during this step (cat .cargo/mutants.toml matches step 8's report).
Per-crate mutant inventory at step 8 (cargo-mutants mutants --list -p <crate> --json | jq length, respecting the widened exclude_globs): smith 63, smith-core 498, smith-ai 268, smith-rpc 29, smith-harness 144, smith-cli 22, smith-alloc 19 — 1043 total. No workspace-wide or full per-crate mutants run has been executed yet (nightly-only, by design); this step did not run one either.
Last scoped result: cargo x mutants -f smith/src/config.rs — 19 mutants, 18 caught, 1 unviable, 0 missed (baseline before step 8's fix was 7 missed).
Coverage summary
cargo x coverage run once during this step (cargo-llvm-cov nextest, both tiers, then --summary-only). 19 e2e tests passed within the coverage run; TOTAL line:
TOTAL 13959 2033 85.44% 1004 199 80.18% 9095 1426 84.32% 0 0 -
Regions 85.44%, functions 80.18%, lines 84.32%, branches not instrumented (0/0, -). Lowest-covered production files: xtask/src/main.rs 17.04% lines (automation, not first-party test target — excluded from mutation scope for the same reason), smith-harness/src/wasm/sealed_wasi.rs 35.98% (WASM sandbox edges), smith/src/provider.rs 41.67%, smith/src/message.rs 66.67%, smith/src/error.rs 64.56% (error-display paths). plugins/echo/src/lib.rs 0% (unbuilt frontend consumer, matches mutation exclusion). This is a map, not a score, per RULES; no threshold is gated on it.
Sleeps remaining and their recorded decision
grep -rn 'sleep' --include='*.rs' across test files, after step 4's audit:
smith-harness/tests/http_stream.rs:179—std::thread::sleep(gap / 3). keep-with-deadline: fixture pacing, explicitly allowed by RULES ("Fixtures may pace input"); the comment at line 178 says so.smith-core/src/bash.rs,smith-core/src/tools.rs,smith-core/src/agent.rs:sleep 30/sleep 5strings are shell scripts run as the subprocess under test (to hold a process open for signal/cancel/timeout tests), not the Rust test waiting on the system under test — out of scope for the sleep-audit law, unchanged since step 4.- No
tokio::time::sleepor polling loop remains anywhere in test code. Every other wait site from step 4's audit (rpc settle, auth expiry, process-exit signal, store repair, tool cancel) uses an injected clock, a channel signal, or a synchronize-then-check pattern with a deadline, per that step's report.
Deferred items
Fuzz targets (deferred until the nightly-toolchain footprint is understood, per the plan's Decisions):
- the frame reader (
smith-core::frame::read_frames, arbitrary byte streams — the property test only covers well-formed round trips, not corruption) EntryContent::decode(arbitrary CBOR bytes, not just re-encoded valid content)- the three SSE decoders in
smith-ai(Anthropic, Gemini, OpenAI transports — arbitrary byte chunks over the wire) smith-rpc'shandle_line(arbitrary JSON-RPC input lines, beyond the one malformed-line e2e case)- constraint:
cargo-fuzzrequires nightly; the plan defers acting on this until that footprint (toolchain pin, CI runner, corpus storage) is scoped as its own decision.
Frontend e2e harness (plan step 11, waits for frontends to land): one harness skeleton per smith-ui consumer — smith-tui (PTY driver via expectrl, screen snapshots, terminal-state checks), an MCP consumer (real MCP client over stdio), a web consumer (browser driver). None of smith-tui, smith-ui, or an MCP/web frontend exist yet as buildable code (0 test-bearing files under smith-tui/src, smith-ui/src; confirmed above). expectrl is declared in [workspace.dependencies] with no member consumer yet — see below.
Root [workspace.dependencies] dev entries with no user after this plan
Checked with grep -rn "^<dep>" --include='Cargo.toml' . (member declarations) and grep -rln "<dep>::" (actual usage in .rs), across every member Cargo.toml:
| dep | declared in a member Cargo.toml? |
used (dep:: in any .rs)? |
verdict |
|---|---|---|---|
| insta | smith-cli, smith-core |
yes (5 snapshot call sites) | keep |
| proptest | smith-core |
yes (9 property tests) | keep |
| criterion | smith |
yes (smith/benches/message_creation.rs) |
keep |
| expectrl | none | no | candidate for the human owner: reserved for the step-11 TUI e2e harness (plan says "PTY via expectrl"); currently dead weight in the root manifest with zero consumers workspace-wide |
| assert_fs | none | no | candidate for the human owner: no test anywhere uses it; assert_cmd + manual tempfile/TempDir cover the current CLI e2e fixtures instead |
| tokio | none | no | candidate for the human owner: marked # dev-only until an explicit runtime decision; no crate has adopted async test infrastructure yet |
cargo x arch (which runs cargo-machete) exits 0 right now, confirming every dependency currently declared in a member manifest is used — there is nothing to remove there. The three flagged above are declared only at the workspace level (root Cargo.toml, human-owned, out of this step's boundary) and have zero consumers; removing or keeping them is the human owner's call, not this step's to make.
Member Cargo.toml dev-dependency audit (this step's own check)
Every dev-dependency actually declared in a member manifest, verified used:
smith:criterion(bench),ciborium(used by session/message tests — not re-verified here, unchanged by this step,cargo-macheteconfirms).smith-cli:archunit,assert_cmd,predicates,insta,smith-core(path, dev). All used per steps 2/3/6's reports andcargo-machete.smith-core:smith-alloc(path),proptest,insta. Used per steps 5/6.smith-harness:serde_json,smith-alloc(path). Used (bounds/http tests).smith-rpc:smith-harness(path),futures. Used.smith-ai,smith-alloc,smith-tui,xtask,smith-ui: no[dev-dependencies]section, or empty.
No member-manifest edits were made: cargo-machete (part of cargo x arch) found nothing unused, and jj diff --stat confirms no Cargo.toml under smith*/ changed.
Criterion bench fix
smith/benches/message_creation.rs: replaced the deprecated criterion::black_box import with std::hint::black_box, keeping the same call sites. cargo bench -p smith --no-run builds clean (verified after the change). Cargo.lock unaffected (cargo check --offline at workspace root ran clean with no lockfile diff — expected, since no manifest changed).
Deviations
- No dev-dependency was removed from any member
Cargo.toml:cargo-machete(already wired intocargo x archsince before this plan) reports zero unused dependencies in every member manifest as they stand. The task's "remove unused dev-dependencies from member manifests" therefore had nothing to act on; the unused entries (expectrl,assert_fs,tokio) exist only in the human-owned root[workspace.dependencies]table, listed above as candidates instead of touched. cargo updatewas not run (as instructed);Cargo.lockhas no diff from this step.