Strong rules.
Read before work.
Obey unless user overrides.
Law
Spec before code.
No production .rs without matching spec coverage.
Cargo is the only build system.
No make, just, root scripts, package scripts, or non-cargo task runners.
No configuration files in the repository root: tool configuration lives under .config/, inside Cargo.toml, or .cargo/ where a tool hardcodes that path; rust-toolchain.toml and lockfiles excepted as toolchain-mandated.
Latest stable Rust, edition 2024.
No numeric MSRV unless release policy creates one.
If layers conflict, higher layer wins.
Fix lower layer.
Do not guess.
Open question: stop and ask.
Build gates
Zero warnings.
Lint scopes: production strict with a curated high-signal allow-list; tests relax assertions and pedantic; guest crates stay strict on their wasm target.
No unwrap/expect/panic/todo/unimplemented in shipped library code.
Tests may use unwrap/expect/panic.
Prefer #[expect(..., reason = "...")] over #[allow].
Proc-macro warning suppression must be narrow and justified on the deriving item or containing module.
#![forbid(unsafe_code)] in library crates except smith-alloc.
smith-cli unsafe only for OS terminal manipulation; smith-alloc unsafe only for GlobalAlloc. Both with safe wrapper, // SAFETY:, and tests.
Public docs required.
fmt, clippy, tests, docs, audit, arch, and architecture-test gates in CI.
Every cargo x command accepts argument passthrough; -p scoping overrides workspace defaults. Lockfile-wide commands (audit, bootstrap, ci) reject scope explicitly.
Build steps run --locked; lockfile drift fails instead of silently updating.
Testing
Every file with test code names the spec it proves (SMH-SPEC-…) in its header; cargo x arch rejects test code without one.
No retries; a flaky test is a bug. Tests never sleep or poll to wait for the system under test; time, cancellation, and readiness are injected or synchronized. Fixtures may pace input; every wait carries a deadline.
Regression first: a bug fix lands with the test that failed before it; a refactor starts with a test that pins the behavior it must keep; a security finding lands with the test that encodes the boundary.
Kinds, each where it is strong:
end to end, core: the real binary through cli, eval, and rpc; real files, real sockets, scripted provider; every user-visible workflow has one; the acceptance gate for correctness.
end to end, frontends: every smith-ui consumer owns a harness fit to its technology (tui: PTY driver with screen snapshots and terminal-state checks; mcp: a real MCP client over stdio; web: a browser driver); the intent vocabulary is the contract and degradation per the intent law is asserted, not assumed; the acceptance gate for usability.
unit and behavior: core crates; one owner per spec sentence; mock StreamFn; isolated dirs.
property: pure functions with algebraic laws (encode and decode, canonicalization, span and index math, masking); proptest with checked-in regression files.
snapshot: rendered frames, RPC lines, traces, help text; insta, reviewed diffs, redacted ids and times.
conformance kits: providers, plugin guests, frontends; fixtures in, vocabulary out.
golden corpus: checked-in sessions; replay and trace byte-identical; regenerated only with a frame-version bump.
bounds: allocation counts and type sizes; exact pins; tighten only.
instruction budgets: callgrind instruction counts per hot path (gungraun), pinned in smith-bench/budgets.toml with a ±2 % band for the platform (toolchain and libc) named in the pins, the CI image; other platforms report without gating; a regression fails, an improvement fails until re-pinned; re-pins are deliberate commits, toolchain bumps the expected trigger.
wall clock: criterion and hyperfine compare two builds on one machine in one session; numbers are never committed and never gate.
mutation: all first-party crates nightly; missed mutants become issues.
coverage: cargo-llvm-cov over every tier; reported per push as a diff, never a percentage gate; uncovered spec sentences are the backlog.
Tiers: cargo x test (unit, property, snapshot, bounds; every commit) → cargo x e2e (binary; every push) → cargo x perf (instruction budgets; every push) → cargo x coverage (report) → nightly (mutants, golden regeneration check).
Evidence bundles: cargo x report collects every tier (JUnit, coverage, last mutation run) into target/report/; the nightly job and a manual report job publish it.
A dev dependency earns its place by a test that uses it.
smith-bench: any workspace crate; nothing depends on it.
xtask: none.
Dependencies
The sanctioned set lives in [workspace.dependencies], grouped by tier; member crates reference entries, never pin versions locally.
Tiers:
vocabulary: ecosystem trait and IR crates; sanctioned once; facade-exempt.
engines: large-surface crates; each owned by exactly one workspace crate; app vocabulary at boundaries; one stack per concern.
leaf utilities: small crates used directly.
build tools: external binaries installed by xtask with --locked; never crate dependencies.
External types never appear in public workspace-crate APIs; vocabulary derives on own types are the exemption.
Prefer std and first-party code; engines only for hard-to-get-right domains.
Pure Rust over system dependencies.
default-features = false unless the feature is the point; enabled features are named.
Application crates define no cargo features; optional behavior lives behind configuration, not compile-time toggles.
No git dependencies; released versions only.
Determinism boundary: nothing on session, replay, or provider paths adds ambient time, randomness, or iteration-order dependence.
Network and filesystem I/O never run on the render thread.
Async runtimes are engine baggage, never a Smith choice: an engine-owner crate may embed one; it never leaks past that crate's boundary.
Concurrency
Core crates (smith, smith-core, smith-ai, smith-ui, smith-rpc) contain no async runtime and no futures in public API.
Core is a pure function over the session arena and one ordered inbound event log; every nondeterministic input (net, fs, clock, plugin output, user input) enters through that log and is recorded at entry.
Threads by role: input, render, core, kernel pool, io. Edges are typed bounded channels; payloads are owned or Arc<immutable>; cancellation is an atomic handle.
Render thread owns the terminal, drains all pending events, then paints once; it never performs I/O and never waits on a producer.
Plugins run as kernels: sync wasmtime, fuel plus epoch, one instance pinned to one thread at a time; effects leave the guest as requests and re-enter as the next call.
Mutex does not appear in library crates.
'static boundaries take snapshots; they never force upstream cloning.
Memory
Lifetime groups are the unit of allocation: process, session, turn, kernel call, frame, io request. Group by lifetime, never by type.
Owner at rest, borrow in flight: structs own; fns take &T, &str, &[T]; Cow only at I/O edges.
Session data is an index arena: contiguous Vec per branch, u32 handles in memory, Uuid at rest.
Large payloads (tool output, provider bodies) are bytes at rest, parsed on demand into the consuming scope.
No clone to satisfy the borrow checker; restructure ownership instead.
Reject drift.
Delete duplication.
Keep source of truth in .system/.
# Rules
Strong rules.
Read before work.
Obey unless user overrides.
## Law
- Spec before code.
- No production `.rs` without matching spec coverage.
- Cargo is the only build system.
- No make, just, root scripts, package scripts, or non-cargo task runners.
- No configuration files in the repository root: tool configuration lives under `.config/`, inside `Cargo.toml`, or `.cargo/` where a tool hardcodes that path; `rust-toolchain.toml` and lockfiles excepted as toolchain-mandated.
- Latest stable Rust, edition 2024.
- No numeric MSRV unless release policy creates one.
- wasmtime plus WIT components for plugins.
- CBOR sessions.
- XDG dirs.
- jj VCS.
- Secret proxy, no keychain dependency.
- StreamFn abstracts providers.
- Plugin wasm runs sandboxed: no ambient WASI, host capabilities only, fuel metered.
- Rust widgets, wasm layout.
## Authority
1. Research evidence.
2. Prototype proof.
3. Specs.
4. Design details.
5. Code.
If layers conflict, higher layer wins.
Fix lower layer.
Do not guess.
Open question: stop and ask.
## Build gates
- Zero warnings.
- Lint scopes: production strict with a curated high-signal allow-list; tests relax assertions and pedantic; guest crates stay strict on their wasm target.
- clippy all/pedantic/perf/style/correctness enforced.
- No unwrap/expect/panic/todo/unimplemented in shipped library code.
- Tests may use unwrap/expect/panic.
- Prefer `#[expect(..., reason = "...")]` over `#[allow]`.
- Proc-macro warning suppression must be narrow and justified on the deriving item or containing module.
- `#![forbid(unsafe_code)]` in library crates except `smith-alloc`.
- `smith-cli` unsafe only for OS terminal manipulation; `smith-alloc` unsafe only for `GlobalAlloc`. Both with safe wrapper, `// SAFETY:`, and tests.
- Public docs required.
- fmt, clippy, tests, docs, audit, arch, and architecture-test gates in CI.
- Every `cargo x` command accepts argument passthrough; `-p` scoping overrides workspace defaults. Lockfile-wide commands (`audit`, `bootstrap`, `ci`) reject scope explicitly.
- Build steps run `--locked`; lockfile drift fails instead of silently updating.
## Testing
- Every file with test code names the spec it proves (`SMH-SPEC-…`) in its header; `cargo x arch` rejects test code without one.
- No retries; a flaky test is a bug. Tests never sleep or poll to wait for the system under test; time, cancellation, and readiness are injected or synchronized. Fixtures may pace input; every wait carries a deadline.
- Regression first: a bug fix lands with the test that failed before it; a refactor starts with a test that pins the behavior it must keep; a security finding lands with the test that encodes the boundary.
- Kinds, each where it is strong:
- end to end, core: the real binary through cli, eval, and rpc; real files, real sockets, scripted provider; every user-visible workflow has one; the acceptance gate for correctness.
- end to end, frontends: every `smith-ui` consumer owns a harness fit to its technology (tui: PTY driver with screen snapshots and terminal-state checks; mcp: a real MCP client over stdio; web: a browser driver); the intent vocabulary is the contract and degradation per the intent law is asserted, not assumed; the acceptance gate for usability.
- unit and behavior: core crates; one owner per spec sentence; mock `StreamFn`; isolated dirs.
- property: pure functions with algebraic laws (encode and decode, canonicalization, span and index math, masking); `proptest` with checked-in regression files.
- snapshot: rendered frames, RPC lines, traces, help text; `insta`, reviewed diffs, redacted ids and times.
- conformance kits: providers, plugin guests, frontends; fixtures in, vocabulary out.
- golden corpus: checked-in sessions; replay and trace byte-identical; regenerated only with a frame-version bump.
- bounds: allocation counts and type sizes; exact pins; tighten only.
- instruction budgets: callgrind instruction counts per hot path (`gungraun`), pinned in `smith-bench/budgets.toml` with a ±2 % band for the platform (toolchain and libc) named in the pins, the CI image; other platforms report without gating; a regression fails, an improvement fails until re-pinned; re-pins are deliberate commits, toolchain bumps the expected trigger.
- wall clock: `criterion` and `hyperfine` compare two builds on one machine in one session; numbers are never committed and never gate.
- mutation: all first-party crates nightly; missed mutants become issues.
- coverage: `cargo-llvm-cov` over every tier; reported per push as a diff, never a percentage gate; uncovered spec sentences are the backlog.
- Tiers: `cargo x test` (unit, property, snapshot, bounds; every commit) → `cargo x e2e` (binary; every push) → `cargo x perf` (instruction budgets; every push) → `cargo x coverage` (report) → nightly (`mutants`, golden regeneration check).
- Evidence bundles: `cargo x report` collects every tier (JUnit, coverage, last mutation run) into `target/report/`; the nightly job and a manual `report` job publish it.
- A dev dependency earns its place by a test that uses it.
## Architecture
```text
smith/ shared types, StreamFn, AgentTool, config, errors
smith-core/ agent loop, sessions, tools, hooks, compaction, cost, trace, replay
smith-ai/ providers, auth, model registry, provider streams, MuxProvider
smith-ui/ abstract UI vocabulary, semantic intents, degradation defaults
smith-tui/ terminal events, widgets, layout primitives, render loop, themes
smith-rpc/ JSON-RPC frontend over standard input and output
smith-harness/ orchestration, wasm plugin host, plugin system, SDK, event bridge, built-ins, help
smith-cli/ binary entry point, CLI, session commands, eval/rpc/replay
smith-alloc/ scope-tagged global allocator, allocation accounting
smith-bench/ benchmarks: instruction budgets, wall-clock A/B; never shipped
xtask/ cargo-only automation
```
Allowed deps:
Library edges; tests may depend on any workspace crate.
- `smith`: none.
- `smith-ui`: only `smith`.
- `smith-core`, `smith-ai`: only `smith`.
- `smith-tui`: `smith`, `smith-ui`.
- `smith-rpc`: `smith`, `smith-core`.
- `smith-harness`: `smith`, `smith-core`, `smith-ai`, `smith-tui`, `smith-ui`.
- `smith-cli`: `smith`, `smith-harness`, `smith-rpc`, `smith-alloc`.
- `smith-alloc`: none.
- `smith-bench`: any workspace crate; nothing depends on it.
- `xtask`: none.
## Dependencies
- The sanctioned set lives in `[workspace.dependencies]`, grouped by tier; member crates reference entries, never pin versions locally.
- Tiers:
- vocabulary: ecosystem trait and IR crates; sanctioned once; facade-exempt.
- engines: large-surface crates; each owned by exactly one workspace crate; app vocabulary at boundaries; one stack per concern.
- leaf utilities: small crates used directly.
- build tools: external binaries installed by xtask with `--locked`; never crate dependencies.
- External types never appear in public workspace-crate APIs; vocabulary derives on own types are the exemption.
- Prefer std and first-party code; engines only for hard-to-get-right domains.
- Pure Rust over system dependencies.
- `default-features = false` unless the feature is the point; enabled features are named.
- Application crates define no cargo features; optional behavior lives behind configuration, not compile-time toggles.
- No git dependencies; released versions only.
- Determinism boundary: nothing on session, replay, or provider paths adds ambient time, randomness, or iteration-order dependence.
- Network and filesystem I/O never run on the render thread.
- Async runtimes are engine baggage, never a Smith choice: an engine-owner crate may embed one; it never leaks past that crate's boundary.
## Concurrency
- Core crates (`smith`, `smith-core`, `smith-ai`, `smith-ui`, `smith-rpc`) contain no async runtime and no futures in public API.
- Core is a pure function over the session arena and one ordered inbound event log; every nondeterministic input (net, fs, clock, plugin output, user input) enters through that log and is recorded at entry.
- Threads by role: input, render, core, kernel pool, io. Edges are typed bounded channels; payloads are owned or `Arc<immutable>`; cancellation is an atomic handle.
- Render thread owns the terminal, drains all pending events, then paints once; it never performs I/O and never waits on a producer.
- Plugins run as kernels: sync wasmtime, fuel plus epoch, one instance pinned to one thread at a time; effects leave the guest as requests and re-enter as the next call.
- `Mutex` does not appear in library crates.
- `'static` boundaries take snapshots; they never force upstream cloning.
## Memory
- Lifetime groups are the unit of allocation: process, session, turn, kernel call, frame, io request. Group by lifetime, never by type.
- Owner at rest, borrow in flight: structs own; fns take `&T`, `&str`, `&[T]`; `Cow` only at I/O edges.
- Session data is an index arena: contiguous `Vec` per branch, `u32` handles in memory, `Uuid` at rest.
- Large payloads (tool output, provider bodies) are bytes at rest, parsed on demand into the consuming scope.
- No clone to satisfy the borrow checker; restructure ownership instead.
- Sharing order: move → `&T` → `Arc<immutable>` → atomic → channel.
- `smith-alloc` attributes every allocation, including engine allocations, to the current lifetime group.
- Ceilings are soft in production (recorded, reported) and hard in tests; hard at runtime only where Smith calls `try_reserve`.
- Hot paths carry allocation-count tests; bounds tighten, never loosen.
- Underlying allocator choice is a benchmark result, never a default.
- Unused dependencies are flagged by `cargo x check` and removed.
- Upgrades are deliberate: `cargo update` in its own commit; lockfile committed.
- Duplicate-version warnings are resolved by upgrades, never by skip-lists.
- Enforcement is mechanical; humans judge principles, never gatekeep entries.
## Research
- Research is evidence, not authority.
- Research docs collect facts, tradeoffs, links, commands, measurements, and examples.
- Keep conclusions as candidates, not mandates.
- External code is untrusted. Do not copy code without license review.
- Do not vendor external repos into this repo.
## Agent boundaries
Agents must not edit without explicit approval:
- `.system/MISSION.md`
- `.system/RULES.md`
- `.system/specs/`
- `Cargo.toml` workspace root
- `.cargo/config.toml`
- `rust-toolchain.toml`
Agents may edit with normal care:
- implementation files,
- tests,
- xtask code,
- benches,
- `.system/research/`, `.system/issues/`, `.system/plans/`.
## Review
Reject drift.
Delete duplication.
Keep source of truth in `.system/`.