--- id: SMH-RESEARCH-4KD6V6IH type: research title: "Concurrency and Memory Shape" --- # Concurrency and Memory Shape Date: 2026-09-26. Status: evidence from code survey and three in-repo spikes; candidates, not mandates. Supersedes the runtime-related parts of SMH-RESEARCH-RSCH0001 and SMH-RESEARCH-RSCH0010 (both predate the wasm plugin law and assume Lua and an async provider boundary). ## Question Which thread, runtime, ownership, and allocation shape gives Smith a render loop that never stalls, plugins that run as isolated kernels, bounded network and file I/O, deterministic replay, and cache-efficient data, while staying small enough that later web, daemon, and p2p frontends attach without reshaping the core? ## Facts ### Code survey at `tztonmyo` (≈11.5k loc host crates) | crate | loc | `.clone()` | `Arc<` | `Box` per call, then clones every `Entry` up the chain. Measured: 2020 allocations, 549 KiB for `fork_at` over 1000 entries (spike C). - `smith-core/src/agent.rs` request build: clones the full message history per provider turn, then masks in place. - `Entry.content` embeds `serde_json::Value`: one heap node per JSON node. - `smith::provider::RecordSpool = Arc>>`: three indirections plus a lock on the recording path. - `smith::provider::StreamFn` returns `BoxStream<'static>`; no consumer awaits it, the harness drives it from a thread. - `smith-harness/src/http.rs` is already blocking: reader thread, `std::sync::mpsc`, `CancelHandle` atomic. - `clippy::needless_pass_by_value` is allowed workspace-wide. ### Runtime dependency roots at `tztonmyo` `tokio` enters only through two engines, never by Smith's choice: ```text tokio 1.53 ├── reqwest 0.13 → hyper 1, hyper-util, tower, tokio-util, tokio-rustls, aws-lc-sys └── wasmtime-wasi 49 (hard dep; feature p2 enables wasmtime/async; sync API is block_on) ``` `wasmtime-wasi` also pulls `rand`, `cap-primitives`, `rustix` (ambient OS surface). `wasmtime-wasi-io` 49 is `no_std` and tokio-free but hard-enables `wasmtime/async` and still leaves clocks and cli unimplemented. The echo guest (`wasm32-wasip2`, wit-bindgen 0.62) imports exactly: `wasi:io/{poll,error,streams}`, `wasi:clocks/monotonic-clock`, `wasi:cli/{stdin,stdout,stderr,environment,exit,terminal-*}` at 0.2.9; no filesystem, random, sockets, or wall-clock. Guest 0.2.9 imports resolve against host 0.2.12 definitions through semver matching. ### Allocator control on stable Rust | mechanism | stable | effort | fit | |---|---|---|---| | `#[global_allocator]` swap or wrapper | yes | one `unsafe impl GlobalAlloc` in a bin or dedicated crate | counting, bounding, attribution of all code including engines | | `allocator_api` (`Vec::new_in`) | no | nightly; `String` has no `_in`; engines allocate globally anyway | rejected | | lifetime arena (`bumpalo`) | yes | `'a` infects async, serde, wasm host, plugin API | rejected for core; scratch only where no reference escapes | | index arena (`Vec` + `u32` handle) | yes | layout discipline only | chosen for session data | Lifetime grouping by type is almost never the logical grouping; Smith's logical groups are process, session, turn, kernel call, frame, io request. ## Spikes All three ran in-repo as sibling changes off `tztonmyo` (shared build cache); each is a candidate for promotion after spec coverage. ### A. Sealed WASI shims replace `wasmtime-wasi` (`qwzlzyqk` / `db1b1af4`) Result: supported. - One `wasmtime::component::bindgen!` call covers the plugin world plus an inline `smith:host` world importing exactly the guest's 12 WASI interfaces; generates sync `Host` traits; `exit` marked trappable. - Shims in `smith-harness/src/wasm/sealed_wasi.rs`: empty environment/arguments/cwd, stdin closed, stdout/stderr accept and discard, terminals `None`, exit traps, monotonic clock returns host-injected value (0 at load), every pollable ready, no `ResourceTable`, no `WasiCtx`. - WIT: `wit/deps/{io,clocks,cli}.wit` copied from wasmtime-wasi 49.0.1 (Apache-2.0 WITH LLVM-exception); `cli.wit` has its `imports`/`command` worlds removed because they reference filesystem, sockets, random; header states the modification. - Gates: `cargo x check|test|lint -p smith-harness` green, 12 tests incl. echo end to end; negative control (drop one import) fails instantiation as expected. - Tree: 298 → 277 crates; `wasmtime feature "async"` present → absent; `cap-primitives`, `rand`, `wasmtime-wasi` gone; `tokio` remains only via reqwest. - Risks: other guests importing random/filesystem/sockets/wall-clock fail at instantiation until further sealed shims land; nothing recorded to the session trace yet; clock frozen; `smith-cli/tests/architecture.rs` still names `wasmtime-wasi`; pre-existing bench failure (`criterion::black_box` deprecated) untouched. ### B. `ureq` 3 HTTP executor (`lnrvruyv` / `057ac857`) Result: supported. - `UreqHttpExecutor` behind the existing `HttpExecutor` trait; shared `admit()` and `stream_body()` pump used by both engines; cancel checked after each read, before send. - `ureq = { version = "3.4", default-features = false, features = ["rustls"] }` → ureq 3.4.2, ureq-proto 0.6.4, ring, webpki-roots. - Loopback SSE test (3 events, 200 ms apart), both engines: first chunk 1.55 ms (ureq) vs 0.84 ms (reqwest); last 401 ms both; cancel → channel closed ≤ 200 ms; body cap honoured at 64 MiB. - Gap timeout (300 ms, server stalls 1 s): body ends at 605 ms. Built as a 40-line `GapTimeout` connector over `ureq::unversioned::transport`, which is semver-unstable across minor releases. ureq has no stable per-read timeout; `timeout_recv_body` is a whole-body budget. - reqwest's blocking client had no per-chunk gap timeout; its per-request `.timeout` acts on send and each body read. - Tree with reqwest removed: 298 → 271 crates (−27 raw, −15 unique normal packages, −45 lockfile entries); removes hyper, tower, tokio-util, hyper-util, tokio-rustls, aws-lc-sys. `reqwest` is referenced only in `smith-harness/src/http.rs`, `smith-harness/Cargo.toml`, root `Cargo.toml`. - TLS trust store changes from platform verifier to bundled Mozilla roots unless `platform-verifier` is enabled. - Redirect, proxy env, and user-agent defaults not compared. ### C. Scope-tagged counting allocator (`onllulxs` / `e124927f`) Result: supported. - New `smith-alloc` library crate (unsafe permitted, `// SAFETY:` comments, 5 unit tests): `Scope` enum (Process, Session, Turn, KernelCall, Frame, IoRequest), RAII `enter`, `snapshot`, `#[global_allocator]` set in `smith-cli`. - Design: `thread_local!` with `const { Cell<[Scope; 16] + len> }` (Copy, no Drop, never allocates, no TLS destructor); per-scope `AtomicU64 × 4` (allocations, bytes, live, peak) Relaxed; `align`-byte header per block carrying the scope tag so `dealloc`/`realloc` charge the allocating scope; guard restores `min(len, previous)`; falls back to Process during TLS teardown. - No locks, no allocation in the alloc path; reentrancy verified by exact unit counts. - Mock eval turn, XDG isolated, three runs identical: Turn 222 allocations / 20 552 B / peak 17 359 B; Process 573 / 134 446 B / peak 70 056 B. Without XDG isolation counts drift because `eval.smh` grows between runs. - `fork_at` over 1000 entries: 2020 allocations, 549 438 B, peak 334 946 B; two consecutive forks equal. - Overhead: whole mock eval process −1 % … +8 % (noise); alloc-heavy microbench +12 % ≈ 12 ns per allocation over ≈ 53 ns; four atomic RMWs per alloc dominate, thread-local counters would cut it. - Risks: header changes memory layout (allocator comparisons must account for it); parallel tests sharing a scope mix counters; peak is lifetime not per-guard; spike-only `--cfg smith_alloc_off` toggle and `SMITH_ALLOC_STATS` stderr dump must go; `smith-alloc` missing from xtask `CRATES`; RULES unsafe law and spec coverage needed. ## Tradeoffs - Async core vs sync core: async buys nothing at Smith's concurrency scale (a handful of provider streams, one bash process, sync grep) and costs `Send + 'static` on every seam, `Arc` infection, and scheduler-dependent ordering inside the determinism boundary. Sync core with thread roles makes replay a pure function over the recorded inbound log. - Runtime crate choice is moot while engines dictate it; the actionable law is confinement (runtime never leaks past the engine-owner crate), not prohibition. p2p (`iroh`) and a web frontend will bring tokio back inside their own crates. - `wasmtime-wasi-io` vs own shims: the former keeps `wasmtime/async` on and covers only io; own shims are ≈ 300 lines, fully sync, and match the spec's "Smith implementations" wording, at the cost of maintaining a hand-stripped `cli.wit`. - `ureq` vs reqwest: −27 crates, no runtime, ring instead of aws-lc; gap timeout depends on an unstable API and trust-store semantics change. reqwest keeps hyper's maturity and platform roots at the cost of the whole tokio stack. - Global counting allocator: ≈ 12 ns per allocation and one header per block; buys attribution of engine allocations to logical scopes and deterministic count tests. Bump-routing scope allocations through the global allocator is unsafe (engine caches, `thread_local!`, lazy statics escape scopes); real frees stay, bump-reset applies only to arenas Smith owns. - Hard ceilings in a global allocator abort on infallible `Vec::push` inside engines; ceilings are therefore soft in production, hard only in tests and where Smith calls `try_reserve`. ## Candidates ### Concurrency shape ```text core = pure fn (session arena, ordered inbound event log) → (session arena', outbound effects) ``` | thread | owns | never | |---|---|---| | input | terminal events → coalesce → ui channel | blocks on core | | render | terminal, frame budget, drains ui channel | I/O, waits on producers | | core | Session arena, agent loop, single consumer of inbound channel | races, ambient time or randomness | | kernel pool | one plugin instance pinned per worker, sync wasmtime, fuel + epoch | blocking host calls inside guest | | io | net + fs on blocking threads, bounded chunks, cancel by atomic | render-thread work | Edges: typed bounded channels; payloads owned or `Arc`; no `Mutex` in library crates; `'static` boundaries take snapshots, never force upstream cloning. Kernel = `(instance, call, inputs) → outputs + effect requests`; effects are dispatched by core to io and re-enter as the next kernel call, so every effect is recorded and replayable and one instance is never on two threads. Render: drain all pending events, then one frame; on-demand with a 16 ms deadline; layout from the wasm layout plugin runs on the kernel pool and returns an owned tree; producers block on full channels, the renderer never does. Later frontends: web (own HTTP server runtime in its crate), daemon (listener plus one core thread per session), p2p (`iroh`, tokio confined to that crate). ### Lifetime groups | group | lives | dies | contents | |---|---|---|---| | process | start | exit | config, model catalog, compiled wasm modules, tool registry | | session | open | close | entry arena, text pool, payload bytes | | turn | user input | provider stop | request buffers, response accumulation, tool-round scratch | | kernel call | host→guest | guest returns | guest linear memory, host conversion buffers | | frame | tick | paint done | layout tree, cell buffer, diff | | io request | send | last chunk | chunk buffers, decoder state | ### Data layout - Session entries: contiguous `Vec` per branch, fixed-size `Entry` with `u32` spans into a per-session text pool and a payload byte arena; `EntryId(Uuid)` stays the at-rest identity, `EntryIdx(u32)` the in-memory handle; ancestry becomes an index chase. - Payloads (tool input, provider bodies): canonical bytes at rest, parsed on demand into turn scope. - Frame: double-buffered `Vec`, never reallocated after the first frame. - Turn scratch: one bump `Vec`; provider request serialized straight into it; `clear()` per turn. - Underlying allocator (`System` vs `mimalloc`) decided by benchmark once count tests exist; header cost of the counter must be excluded from that comparison. ### Promotion order 1. Spike A (sealed WASI) after spec coverage of sealed semantics and trace recording; fix architecture test naming. 2. Spike C (`smith-alloc`) after RULES unsafe exception and spec coverage; drop spike toggles; add to xtask `CRATES`; switch counters to thread-local. 3. Spike B (`ureq`) after deciding trust-store policy and pinning `~3.4`; delete reqwest engine and workspace entry. 4. Hot-spot refactors with count bounds: `ancestry` → request build → spool → `Value` payloads. ## Promotion results (2026-09-26, stack `tlszyzws..wqptrptw`) All three spikes promoted after spec text landed (SMH-SPEC-WASM0001 Sealed WASI; SMH-SPEC-SPEC0001 network transport, memory accounting, index arena). | measure | before | after | |---|---|---| | crates in normal graph | 298 | 210 | | lockfile packages | 370 | 294 | | async runtime in graph | tokio via reqwest and wasmtime-wasi | none | | `Mutex` in library production code | 5 sites | 0, enforced by `cargo x arch` | | `fork_at(1000)` allocations | 2020 | 3 | | `Session::from_frames(1000)` allocations | unmeasured | 5, equal at 100 entries | | mock eval turn allocations | 222 | 42 | What moved each number: - `ancestry` as branch prefix: 2020 → 2002 (the floor of owned `Entry` clones). - Session arena (SMH-PLAN-82POBOKB): `Entry` is a `Copy` record over one `Vec` arena, content is canonical CBOR bytes decoded on demand, frames embed content as one byte string (`VERSION_ENTRY = 2`): 2002 → 3, load constant, turn 220 → 208. - `Arc<[ToolMetadata]>` registry snapshot on `ProviderRequest.tools` instead of a per-request sorted clone of every schema: 208 → 42. - Provider records over bounded `mpsc`, single-pass masked request build: 222 → 220 in-process (the first turn per process pays +1 allocation of lazy global init; bounds are steady state after a warm-up). - HTTP body channel bounded (16 chunks) and terminated by an explicit `HttpBodyEnd`; RPC worker owns the agent, reader owns responses, two atomics plus channels. Decisions taken on the way: bundled Mozilla roots for TLS (`platform-verifier` deferred); `record_channel` bound 64 with a documented per-round invariant; `smith-alloc` counters stay global atomics (≈ 12 ns per allocation) until a hot path shows the cost; per-frame content `Vec` on load accepted until the reader hands slices. Remaining floor of the mock turn (42): request encoding, stream event vectors, two entry encodes and frame writes; next candidates are a slice-handing frame reader, a borrowed `EntryContent` view at record time, and a per-session text pool. ## Open questions - Kernel pool sizing: lazy pool capped at cores vs one thread per plugin. - Render tick: fixed 60 Hz vs on-demand with deadline. - Sealed WASI breadth: which further imports (random seed, wall-clock, filesystem) get Smith substitutes before a non-trivial guest std program runs; `monotonic_now` is still never injected. - Per-turn peak: reset or epoch API on `smith-alloc`; thread-local counters. - Trust store: revisit `platform-verifier` when a corporate proxy case appears. ## Sources - Code and lockfile at `tztonmyo` (`dd069f92`); spike changes `qwzlzyqk`, `lnrvruyv`, `onllulxs`; promoted stack `tlszyzws` through `wqptrptw`. - Registry sources: `wasmtime-wasi 49.0.1`, `wasmtime-wasi-io 49.0.1`, `ureq 3.4.2`, `ureq-proto 0.6.4`. - SMH-SPEC-WASM0001 (sealed WASI law, single-threaded instances), SMH-SPEC-SPEC0001 (bounded frame allocation). - SMH-RESEARCH-RSCH0001 (arena patterns), SMH-RESEARCH-RSCH0010 (session size distribution: p99 ≈ 9 MiB JSONL, compaction near 250k tokens).