# Rules Strong rules. Read before work. Obey unless user overrides. ## Law - Spec before code. - No production `.rs` without matching spec coverage. - Cargo is the only build system. - No make, just, root scripts, package scripts, or non-cargo task runners. - No configuration files in the repository root: tool configuration lives under `.config/`, inside `Cargo.toml`, or `.cargo/` where a tool hardcodes that path; `rust-toolchain.toml` and lockfiles excepted as toolchain-mandated. - Latest stable Rust, edition 2024. - No numeric MSRV unless release policy creates one. - wasmtime plus WIT components for plugins. - CBOR sessions. - XDG dirs. - jj VCS. - Secret proxy, no keychain dependency. - StreamFn abstracts providers. - Plugin wasm runs sandboxed: no ambient WASI, host capabilities only, fuel metered. - Rust widgets, wasm layout. ## Authority 1. Research evidence. 2. Prototype proof. 3. Specs. 4. Design details. 5. Code. If layers conflict, higher layer wins. Fix lower layer. Do not guess. Open question: stop and ask. ## Build gates - Zero warnings. - Lint scopes: production strict with a curated high-signal allow-list; tests relax assertions and pedantic; guest crates stay strict on their wasm target. - clippy all/pedantic/perf/style/correctness enforced. - No unwrap/expect/panic/todo/unimplemented in shipped library code. - Tests may use unwrap/expect/panic. - Prefer `#[expect(..., reason = "...")]` over `#[allow]`. - Proc-macro warning suppression must be narrow and justified on the deriving item or containing module. - `#![forbid(unsafe_code)]` in library crates except `smith-alloc`. - `smith-cli` unsafe only for OS terminal manipulation; `smith-alloc` unsafe only for `GlobalAlloc`. Both with safe wrapper, `// SAFETY:`, and tests. - Public docs required. - fmt, clippy, tests, docs, audit, arch, and architecture-test gates in CI. - Every `cargo x` command accepts argument passthrough; `-p` scoping overrides workspace defaults. Lockfile-wide commands (`audit`, `bootstrap`, `ci`) reject scope explicitly. - Build steps run `--locked`; lockfile drift fails instead of silently updating. ## Testing - Every file with test code names the spec it proves (`SMH-SPEC-…`) in its header; `cargo x arch` rejects test code without one. - No retries; a flaky test is a bug. Tests never sleep or poll to wait for the system under test; time, cancellation, and readiness are injected or synchronized. Fixtures may pace input; every wait carries a deadline. - Regression first: a bug fix lands with the test that failed before it; a refactor starts with a test that pins the behavior it must keep; a security finding lands with the test that encodes the boundary. - Kinds, each where it is strong: - end to end, core: the real binary through cli, eval, and rpc; real files, real sockets, scripted provider; every user-visible workflow has one; the acceptance gate for correctness. - end to end, frontends: every `smith-ui` consumer owns a harness fit to its technology (tui: PTY driver with screen snapshots and terminal-state checks; mcp: a real MCP client over stdio; web: a browser driver); the intent vocabulary is the contract and degradation per the intent law is asserted, not assumed; the acceptance gate for usability. - unit and behavior: core crates; one owner per spec sentence; mock `StreamFn`; isolated dirs. - property: pure functions with algebraic laws (encode and decode, canonicalization, span and index math, masking); `proptest` with checked-in regression files. - snapshot: rendered frames, RPC lines, traces, help text; `insta`, reviewed diffs, redacted ids and times. - conformance kits: providers, plugin guests, frontends; fixtures in, vocabulary out. - golden corpus: checked-in sessions; replay and trace byte-identical; regenerated only with a frame-version bump. - bounds: allocation counts and type sizes; exact pins; tighten only. - instruction budgets: callgrind instruction counts per hot path (`gungraun`), pinned in `smith-bench/budgets.toml` with a ±2 % band for the platform (toolchain and libc) named in the pins, the CI image; other platforms report without gating; a regression fails, an improvement fails until re-pinned; re-pins are deliberate commits, toolchain bumps the expected trigger. - wall clock: `criterion` and `hyperfine` compare two builds on one machine in one session; numbers are never committed and never gate. - mutation: all first-party crates nightly; missed mutants become issues. - coverage: `cargo-llvm-cov` over every tier; reported per push as a diff, never a percentage gate; uncovered spec sentences are the backlog. - Tiers: `cargo x test` (unit, property, snapshot, bounds; every commit) → `cargo x e2e` (binary; every push) → `cargo x perf` (instruction budgets; every push) → `cargo x coverage` (report) → nightly (`mutants`, golden regeneration check). - Evidence bundles: `cargo x report` collects every tier (JUnit, coverage, last mutation run) into `target/report/`; the nightly job and a manual `report` job publish it. - A dev dependency earns its place by a test that uses it. ## Architecture ```text smith/ shared types, StreamFn, AgentTool, config, errors smith-core/ agent loop, sessions, tools, hooks, compaction, cost, trace, replay smith-ai/ providers, auth, model registry, provider streams, MuxProvider smith-ui/ abstract UI vocabulary, semantic intents, degradation defaults smith-tui/ terminal events, widgets, layout primitives, render loop, themes smith-rpc/ JSON-RPC frontend over standard input and output smith-harness/ orchestration, wasm plugin host, plugin system, SDK, event bridge, built-ins, help smith-cli/ binary entry point, CLI, session commands, eval/rpc/replay smith-alloc/ scope-tagged global allocator, allocation accounting smith-bench/ benchmarks: instruction budgets, wall-clock A/B; never shipped xtask/ cargo-only automation ``` Allowed deps: Library edges; tests may depend on any workspace crate. - `smith`: none. - `smith-ui`: only `smith`. - `smith-core`, `smith-ai`: only `smith`. - `smith-tui`: `smith`, `smith-ui`. - `smith-rpc`: `smith`, `smith-core`. - `smith-harness`: `smith`, `smith-core`, `smith-ai`, `smith-tui`, `smith-ui`. - `smith-cli`: `smith`, `smith-harness`, `smith-rpc`, `smith-alloc`. - `smith-alloc`: none. - `smith-bench`: any workspace crate; nothing depends on it. - `xtask`: none. ## Dependencies - The sanctioned set lives in `[workspace.dependencies]`, grouped by tier; member crates reference entries, never pin versions locally. - Tiers: - vocabulary: ecosystem trait and IR crates; sanctioned once; facade-exempt. - engines: large-surface crates; each owned by exactly one workspace crate; app vocabulary at boundaries; one stack per concern. - leaf utilities: small crates used directly. - build tools: external binaries installed by xtask with `--locked`; never crate dependencies. - External types never appear in public workspace-crate APIs; vocabulary derives on own types are the exemption. - Prefer std and first-party code; engines only for hard-to-get-right domains. - Pure Rust over system dependencies. - `default-features = false` unless the feature is the point; enabled features are named. - Application crates define no cargo features; optional behavior lives behind configuration, not compile-time toggles. - No git dependencies; released versions only. - Determinism boundary: nothing on session, replay, or provider paths adds ambient time, randomness, or iteration-order dependence. - Network and filesystem I/O never run on the render thread. - Async runtimes are engine baggage, never a Smith choice: an engine-owner crate may embed one; it never leaks past that crate's boundary. ## Concurrency - Core crates (`smith`, `smith-core`, `smith-ai`, `smith-ui`, `smith-rpc`) contain no async runtime and no futures in public API. - Core is a pure function over the session arena and one ordered inbound event log; every nondeterministic input (net, fs, clock, plugin output, user input) enters through that log and is recorded at entry. - Threads by role: input, render, core, kernel pool, io. Edges are typed bounded channels; payloads are owned or `Arc`; cancellation is an atomic handle. - Render thread owns the terminal, drains all pending events, then paints once; it never performs I/O and never waits on a producer. - Plugins run as kernels: sync wasmtime, fuel plus epoch, one instance pinned to one thread at a time; effects leave the guest as requests and re-enter as the next call. - `Mutex` does not appear in library crates. - `'static` boundaries take snapshots; they never force upstream cloning. ## Memory - Lifetime groups are the unit of allocation: process, session, turn, kernel call, frame, io request. Group by lifetime, never by type. - Owner at rest, borrow in flight: structs own; fns take `&T`, `&str`, `&[T]`; `Cow` only at I/O edges. - Session data is an index arena: contiguous `Vec` per branch, `u32` handles in memory, `Uuid` at rest. - Large payloads (tool output, provider bodies) are bytes at rest, parsed on demand into the consuming scope. - No clone to satisfy the borrow checker; restructure ownership instead. - Sharing order: move → `&T` → `Arc` → atomic → channel. - `smith-alloc` attributes every allocation, including engine allocations, to the current lifetime group. - Ceilings are soft in production (recorded, reported) and hard in tests; hard at runtime only where Smith calls `try_reserve`. - Hot paths carry allocation-count tests; bounds tighten, never loosen. - Underlying allocator choice is a benchmark result, never a default. - Unused dependencies are flagged by `cargo x check` and removed. - Upgrades are deliberate: `cargo update` in its own commit; lockfile committed. - Duplicate-version warnings are resolved by upgrades, never by skip-lists. - Enforcement is mechanical; humans judge principles, never gatekeep entries. ## Research - Research is evidence, not authority. - Research docs collect facts, tradeoffs, links, commands, measurements, and examples. - Keep conclusions as candidates, not mandates. - External code is untrusted. Do not copy code without license review. - Do not vendor external repos into this repo. ## Agent boundaries Agents must not edit without explicit approval: - `.system/MISSION.md` - `.system/RULES.md` - `.system/specs/` - `Cargo.toml` workspace root - `.cargo/config.toml` - `rust-toolchain.toml` Agents may edit with normal care: - implementation files, - tests, - xtask code, - benches, - `.system/research/`, `.system/issues/`, `.system/plans/`. ## Review Reject drift. Delete duplication. Keep source of truth in `.system/`.