Findings behind the plan: retries masked timing races; 26 of 30 test-bearing files cite no spec; mutation scope covered two crates and missed seven boundary comparisons in config.rs; insta, proptest, expectrl, assert_fs, tokio were declared and unused; smith_core::trace::replay had no corpus.
Decisions
End to end is the acceptance gate: core e2e drives the built binary through cli, eval, and rpc; each frontend owns a harness fit to its technology.
Retries are zero everywhere; waits are injected or synchronized with deadlines.
Every test file cites its spec; cargo x arch enforces it.
Property, snapshot, golden, bounds, mutation, coverage each have one home and one command; nothing is wired twice.
Fuzzing is deferred until its nightly-toolchain footprint is understood; decoders of untrusted bytes are listed as future targets.
Order
Law and wiring: RULES ## Testing; nextest tiers (default excludes e2e_*, e2e selects them, mutants unchanged, all retries = 0); cargo x e2e, cargo x coverage (cargo-llvm-cov, lcov artifact, summary); CI: push runs ci and e2e; nightly runs mutants and coverage.
Core e2e in smith-cli/tests/e2e_*.rs: eval mock turn, eval --json, rpc over stdin/stdout (ping, prompt, abort, wait, malformed line, shutdown), auth round trip, session file survives a second run and replays its trace.
Spec-reference gate in cargo x arch; headers on every test-bearing file.
Sleep audit: rpc poll loops become settle signals with deadlines; remaining sites reviewed one by one.
Property tests: frame encode and decode, canonical_json idempotence and order independence, Span arithmetic, EntryContent and EntryContentRef byte equality, secret mask and unmask; regression files checked in.
Snapshots with insta: eval --json, rpc line sequences, build_trace, --help; ids and timestamps redacted.
Golden corpus: smith-core/tests/golden/*.smh plus *.trace.json; byte-identical replay; SMITH_GOLDEN=regenerate with a frame-version bump only.
Mutation: scope every first-party crate; fix the seven config.rs misses; nightly job.
Coverage and reports: cargo llvm-cov nextest across tiers; cargo x report bundles JUnit per tier, lcov plus HTML coverage, and the last mutation run into target/report/ with an index.md; nightly and a manual report job publish the bundle; no threshold.
Research: supersede RSCH0013 with the measured state; remove dev dependencies no test uses; fix the criterion bench.
Frontend e2e stubs: one harness skeleton per smith-ui consumer, filled when the frontend lands (tui first, PTY via expectrl).
Interfaces
xtask: e2e, coverage, report; arch gains the spec-reference check.
.ci/: check.kdl (push), e2e.kdl (push), nightly.kdl (schedule, publishes the report), report.kdl (manual, publishes the report).
smith-cli/tests/e2e_*.rs: the core e2e suite.
smith-core/tests/golden/: the replay corpus.
Risks
e2e tests spawn the binary; cold builds dominate CI time; cargo x e2e shares the target dir with ci.
Snapshot review needs discipline: a blind INSTA_UPDATE=always turns the gate into noise; snapshots are reviewed in the diff.
Golden regeneration is a deliberate, versioned act; an accidental regen hides a determinism break.
Coverage numbers vary with e2e process boundaries; the report is a map, not a score.
Exit
Every step green under cargo x check, cargo x e2e, cargo x arch; CI jobs observed running on the public repository.
---
id: SMH-PLAN-DFOG846Z
type: plan
title: "Testing Methodology"
spec: SMH-SPEC-SPEC0001
status: approved
depends_on: [SMH-PLAN-CORE0002]
---
# Testing Methodology
## Entry
- RULES `## Testing` names the kinds, the tiers, and the regression-first law.
- 191 tests, ≈ 4 s, nextest process-per-test; conformance kit for providers; allocation bounds pinned.
- Findings behind the plan: retries masked timing races; 26 of 30 test-bearing files cite no spec; mutation scope covered two crates and missed seven boundary comparisons in `config.rs`; `insta`, `proptest`, `expectrl`, `assert_fs`, `tokio` were declared and unused; `smith_core::trace::replay` had no corpus.
## Decisions
- End to end is the acceptance gate: core e2e drives the built binary through cli, eval, and rpc; each frontend owns a harness fit to its technology.
- Retries are zero everywhere; waits are injected or synchronized with deadlines.
- Every test file cites its spec; `cargo x arch` enforces it.
- Property, snapshot, golden, bounds, mutation, coverage each have one home and one command; nothing is wired twice.
- Fuzzing is deferred until its nightly-toolchain footprint is understood; decoders of untrusted bytes are listed as future targets.
## Order
1. Law and wiring: RULES `## Testing`; nextest tiers (`default` excludes `e2e_*`, `e2e` selects them, `mutants` unchanged, all `retries = 0`); `cargo x e2e`, `cargo x coverage` (`cargo-llvm-cov`, lcov artifact, summary); CI: push runs `ci` and `e2e`; nightly runs `mutants` and `coverage`.
2. Core e2e in `smith-cli/tests/e2e_*.rs`: eval mock turn, eval `--json`, rpc over stdin/stdout (ping, prompt, abort, wait, malformed line, shutdown), auth round trip, session file survives a second run and replays its trace.
3. Spec-reference gate in `cargo x arch`; headers on every test-bearing file.
4. Sleep audit: rpc poll loops become settle signals with deadlines; remaining sites reviewed one by one.
5. Property tests: frame encode and decode, `canonical_json` idempotence and order independence, `Span` arithmetic, `EntryContent` and `EntryContentRef` byte equality, secret mask and unmask; regression files checked in.
6. Snapshots with `insta`: `eval --json`, rpc line sequences, `build_trace`, `--help`; ids and timestamps redacted.
7. Golden corpus: `smith-core/tests/golden/*.smh` plus `*.trace.json`; byte-identical replay; `SMITH_GOLDEN=regenerate` with a frame-version bump only.
8. Mutation: scope every first-party crate; fix the seven `config.rs` misses; nightly job.
9. Coverage and reports: `cargo llvm-cov nextest` across tiers; `cargo x report` bundles JUnit per tier, lcov plus HTML coverage, and the last mutation run into `target/report/` with an `index.md`; nightly and a manual `report` job publish the bundle; no threshold.
10. Research: supersede RSCH0013 with the measured state; remove dev dependencies no test uses; fix the criterion bench.
11. Frontend e2e stubs: one harness skeleton per `smith-ui` consumer, filled when the frontend lands (tui first, PTY via `expectrl`).
## Interfaces
- `xtask`: `e2e`, `coverage`, `report`; `arch` gains the spec-reference check.
- `.config/nextest.toml`: profiles `default`, `e2e`, `mutants`.
- `.ci/`: `check.kdl` (push), `e2e.kdl` (push), `nightly.kdl` (schedule, publishes the report), `report.kdl` (manual, publishes the report).
- `smith-cli/tests/e2e_*.rs`: the core e2e suite.
- `smith-core/tests/golden/`: the replay corpus.
## Risks
- e2e tests spawn the binary; cold builds dominate CI time; `cargo x e2e` shares the target dir with `ci`.
- Snapshot review needs discipline: a blind `INSTA_UPDATE=always` turns the gate into noise; snapshots are reviewed in the diff.
- Golden regeneration is a deliberate, versioned act; an accidental regen hides a determinism break.
- Coverage numbers vary with e2e process boundaries; the report is a map, not a score.
## Exit
- Every step green under `cargo x check`, `cargo x e2e`, `cargo x arch`; CI jobs observed running on the public repository.