Luigit
repositories / smith

smith

There are many coding harnesses - but this one is fast

owned by admin

.system/research/SMH-RESEARCH-4KD6V6IH-concurrency-and-memory-shape/index.md

Raw
Rendered preview

id: SMH-RESEARCH-4KD6V6IH type: research title: "Concurrency and Memory Shape"

Concurrency and Memory Shape

Date: 2026-09-26. Status: evidence from code survey and three in-repo spikes; candidates, not mandates. Supersedes the runtime-related parts of SMH-RESEARCH-RSCH0001 and SMH-RESEARCH-RSCH0010 (both predate the wasm plugin law and assume Lua and an async provider boundary).

Question

Which thread, runtime, ownership, and allocation shape gives Smith a render loop that never stalls, plugins that run as isolated kernels, bounded network and file I/O, deterministic replay, and cache-efficient data, while staying small enough that later web, daemon, and p2p frontends attach without reshaping the core?

Facts

Code survey at tztonmyo (≈11.5k loc host crates)

crate loc .clone() Arc< Box<dyn lifetimes Cow async fn
smith 1275 3 4 0 1 0 0
smith-core 5201 68 4 0 0 0 0
smith-ai 3356 15 8 2 0 0 0
smith-rpc 487 9 1 0 0 0 0
smith-harness 678 7 2 0 1 0 0
smith-cli 451 2 0 0 0 0 0

Observed hot spots:

  • smith-core/src/session.rs ancestry: builds a fresh HashMap<EntryId, &Entry> per call, then clones every Entry up the chain. Measured: 2020 allocations, 549 KiB for fork_at over 1000 entries (spike C).
  • smith-core/src/agent.rs request build: clones the full message history per provider turn, then masks in place.
  • Entry.content embeds serde_json::Value: one heap node per JSON node.
  • smith::provider::RecordSpool = Arc<Mutex<Vec<Value>>>: three indirections plus a lock on the recording path.
  • smith::provider::StreamFn returns BoxStream<'static>; no consumer awaits it, the harness drives it from a thread.
  • smith-harness/src/http.rs is already blocking: reader thread, std::sync::mpsc, CancelHandle atomic.
  • clippy::needless_pass_by_value is allowed workspace-wide.

Runtime dependency roots at tztonmyo

tokio enters only through two engines, never by Smith's choice:

tokio 1.53
├── reqwest 0.13 → hyper 1, hyper-util, tower, tokio-util, tokio-rustls, aws-lc-sys
└── wasmtime-wasi 49 (hard dep; feature p2 enables wasmtime/async; sync API is block_on)

wasmtime-wasi also pulls rand, cap-primitives, rustix (ambient OS surface). wasmtime-wasi-io 49 is no_std and tokio-free but hard-enables wasmtime/async and still leaves clocks and cli unimplemented.

The echo guest (wasm32-wasip2, wit-bindgen 0.62) imports exactly: wasi:io/{poll,error,streams}, wasi:clocks/monotonic-clock, wasi:cli/{stdin,stdout,stderr,environment,exit,terminal-*} at 0.2.9; no filesystem, random, sockets, or wall-clock. Guest 0.2.9 imports resolve against host 0.2.12 definitions through semver matching.

Allocator control on stable Rust

mechanism stable effort fit
#[global_allocator] swap or wrapper yes one unsafe impl GlobalAlloc in a bin or dedicated crate counting, bounding, attribution of all code including engines
allocator_api (Vec::new_in) no nightly; String has no _in; engines allocate globally anyway rejected
lifetime arena (bumpalo) yes 'a infects async, serde, wasm host, plugin API rejected for core; scratch only where no reference escapes
index arena (Vec<T> + u32 handle) yes layout discipline only chosen for session data

Lifetime grouping by type is almost never the logical grouping; Smith's logical groups are process, session, turn, kernel call, frame, io request.

Spikes

All three ran in-repo as sibling changes off tztonmyo (shared build cache); each is a candidate for promotion after spec coverage.

A. Sealed WASI shims replace wasmtime-wasi (qwzlzyqk / db1b1af4)

Result: supported.

  • One wasmtime::component::bindgen! call covers the plugin world plus an inline smith:host world importing exactly the guest's 12 WASI interfaces; generates sync Host traits; exit marked trappable.
  • Shims in smith-harness/src/wasm/sealed_wasi.rs: empty environment/arguments/cwd, stdin closed, stdout/stderr accept and discard, terminals None, exit traps, monotonic clock returns host-injected value (0 at load), every pollable ready, no ResourceTable, no WasiCtx.
  • WIT: wit/deps/{io,clocks,cli}.wit copied from wasmtime-wasi 49.0.1 (Apache-2.0 WITH LLVM-exception); cli.wit has its imports/command worlds removed because they reference filesystem, sockets, random; header states the modification.
  • Gates: cargo x check|test|lint -p smith-harness green, 12 tests incl. echo end to end; negative control (drop one import) fails instantiation as expected.
  • Tree: 298 → 277 crates; wasmtime feature "async" present → absent; cap-primitives, rand, wasmtime-wasi gone; tokio remains only via reqwest.
  • Risks: other guests importing random/filesystem/sockets/wall-clock fail at instantiation until further sealed shims land; nothing recorded to the session trace yet; clock frozen; smith-cli/tests/architecture.rs still names wasmtime-wasi; pre-existing bench failure (criterion::black_box deprecated) untouched.

B. ureq 3 HTTP executor (lnrvruyv / 057ac857)

Result: supported.

  • UreqHttpExecutor behind the existing HttpExecutor trait; shared admit() and stream_body() pump used by both engines; cancel checked after each read, before send.
  • ureq = { version = "3.4", default-features = false, features = ["rustls"] } → ureq 3.4.2, ureq-proto 0.6.4, ring, webpki-roots.
  • Loopback SSE test (3 events, 200 ms apart), both engines: first chunk 1.55 ms (ureq) vs 0.84 ms (reqwest); last 401 ms both; cancel → channel closed ≤ 200 ms; body cap honoured at 64 MiB.
  • Gap timeout (300 ms, server stalls 1 s): body ends at 605 ms. Built as a 40-line GapTimeout connector over ureq::unversioned::transport, which is semver-unstable across minor releases. ureq has no stable per-read timeout; timeout_recv_body is a whole-body budget.
  • reqwest's blocking client had no per-chunk gap timeout; its per-request .timeout acts on send and each body read.
  • Tree with reqwest removed: 298 → 271 crates (−27 raw, −15 unique normal packages, −45 lockfile entries); removes hyper, tower, tokio-util, hyper-util, tokio-rustls, aws-lc-sys. reqwest is referenced only in smith-harness/src/http.rs, smith-harness/Cargo.toml, root Cargo.toml.
  • TLS trust store changes from platform verifier to bundled Mozilla roots unless platform-verifier is enabled.
  • Redirect, proxy env, and user-agent defaults not compared.

C. Scope-tagged counting allocator (onllulxs / e124927f)

Result: supported.

  • New smith-alloc library crate (unsafe permitted, // SAFETY: comments, 5 unit tests): Scope enum (Process, Session, Turn, KernelCall, Frame, IoRequest), RAII enter, snapshot, #[global_allocator] set in smith-cli.
  • Design: thread_local! with const { Cell<[Scope; 16] + len> } (Copy, no Drop, never allocates, no TLS destructor); per-scope AtomicU64 × 4 (allocations, bytes, live, peak) Relaxed; align-byte header per block carrying the scope tag so dealloc/realloc charge the allocating scope; guard restores min(len, previous); falls back to Process during TLS teardown.
  • No locks, no allocation in the alloc path; reentrancy verified by exact unit counts.
  • Mock eval turn, XDG isolated, three runs identical: Turn 222 allocations / 20 552 B / peak 17 359 B; Process 573 / 134 446 B / peak 70 056 B. Without XDG isolation counts drift because eval.smh grows between runs.
  • fork_at over 1000 entries: 2020 allocations, 549 438 B, peak 334 946 B; two consecutive forks equal.
  • Overhead: whole mock eval process −1 % … +8 % (noise); alloc-heavy microbench +12 % ≈ 12 ns per allocation over ≈ 53 ns; four atomic RMWs per alloc dominate, thread-local counters would cut it.
  • Risks: header changes memory layout (allocator comparisons must account for it); parallel tests sharing a scope mix counters; peak is lifetime not per-guard; spike-only --cfg smith_alloc_off toggle and SMITH_ALLOC_STATS stderr dump must go; smith-alloc missing from xtask CRATES; RULES unsafe law and spec coverage needed.

Tradeoffs

  • Async core vs sync core: async buys nothing at Smith's concurrency scale (a handful of provider streams, one bash process, sync grep) and costs Send + 'static on every seam, Arc infection, and scheduler-dependent ordering inside the determinism boundary. Sync core with thread roles makes replay a pure function over the recorded inbound log.
  • Runtime crate choice is moot while engines dictate it; the actionable law is confinement (runtime never leaks past the engine-owner crate), not prohibition. p2p (iroh) and a web frontend will bring tokio back inside their own crates.
  • wasmtime-wasi-io vs own shims: the former keeps wasmtime/async on and covers only io; own shims are ≈ 300 lines, fully sync, and match the spec's "Smith implementations" wording, at the cost of maintaining a hand-stripped cli.wit.
  • ureq vs reqwest: −27 crates, no runtime, ring instead of aws-lc; gap timeout depends on an unstable API and trust-store semantics change. reqwest keeps hyper's maturity and platform roots at the cost of the whole tokio stack.
  • Global counting allocator: ≈ 12 ns per allocation and one header per block; buys attribution of engine allocations to logical scopes and deterministic count tests. Bump-routing scope allocations through the global allocator is unsafe (engine caches, thread_local!, lazy statics escape scopes); real frees stay, bump-reset applies only to arenas Smith owns.
  • Hard ceilings in a global allocator abort on infallible Vec::push inside engines; ceilings are therefore soft in production, hard only in tests and where Smith calls try_reserve.

Candidates

Concurrency shape

core = pure fn (session arena, ordered inbound event log) → (session arena', outbound effects)
thread owns never
input terminal events → coalesce → ui channel blocks on core
render terminal, frame budget, drains ui channel I/O, waits on producers
core Session arena, agent loop, single consumer of inbound channel races, ambient time or randomness
kernel pool one plugin instance pinned per worker, sync wasmtime, fuel + epoch blocking host calls inside guest
io net + fs on blocking threads, bounded chunks, cancel by atomic render-thread work

Edges: typed bounded channels; payloads owned or Arc<immutable>; no Mutex in library crates; 'static boundaries take snapshots, never force upstream cloning.

Kernel = (instance, call, inputs) → outputs + effect requests; effects are dispatched by core to io and re-enter as the next kernel call, so every effect is recorded and replayable and one instance is never on two threads.

Render: drain all pending events, then one frame; on-demand with a 16 ms deadline; layout from the wasm layout plugin runs on the kernel pool and returns an owned tree; producers block on full channels, the renderer never does.

Later frontends: web (own HTTP server runtime in its crate), daemon (listener plus one core thread per session), p2p (iroh, tokio confined to that crate).

Lifetime groups

group lives dies contents
process start exit config, model catalog, compiled wasm modules, tool registry
session open close entry arena, text pool, payload bytes
turn user input provider stop request buffers, response accumulation, tool-round scratch
kernel call host→guest guest returns guest linear memory, host conversion buffers
frame tick paint done layout tree, cell buffer, diff
io request send last chunk chunk buffers, decoder state

Data layout

  • Session entries: contiguous Vec<Entry> per branch, fixed-size Entry with u32 spans into a per-session text pool and a payload byte arena; EntryId(Uuid) stays the at-rest identity, EntryIdx(u32) the in-memory handle; ancestry becomes an index chase.
  • Payloads (tool input, provider bodies): canonical bytes at rest, parsed on demand into turn scope.
  • Frame: double-buffered Vec<Cell>, never reallocated after the first frame.
  • Turn scratch: one bump Vec<u8>; provider request serialized straight into it; clear() per turn.
  • Underlying allocator (System vs mimalloc) decided by benchmark once count tests exist; header cost of the counter must be excluded from that comparison.

Promotion order

  1. Spike A (sealed WASI) after spec coverage of sealed semantics and trace recording; fix architecture test naming.
  2. Spike C (smith-alloc) after RULES unsafe exception and spec coverage; drop spike toggles; add to xtask CRATES; switch counters to thread-local.
  3. Spike B (ureq) after deciding trust-store policy and pinning ~3.4; delete reqwest engine and workspace entry.
  4. Hot-spot refactors with count bounds: ancestry → request build → spool → Value payloads.

Promotion results (2026-09-26, stack tlszyzws..wqptrptw)

All three spikes promoted after spec text landed (SMH-SPEC-WASM0001 Sealed WASI; SMH-SPEC-SPEC0001 network transport, memory accounting, index arena).

measure before after
crates in normal graph 298 210
lockfile packages 370 294
async runtime in graph tokio via reqwest and wasmtime-wasi none
Mutex in library production code 5 sites 0, enforced by cargo x arch
fork_at(1000) allocations 2020 3
Session::from_frames(1000) allocations unmeasured 5, equal at 100 entries
mock eval turn allocations 222 42

What moved each number:

  • ancestry as branch prefix: 2020 → 2002 (the floor of owned Entry clones).
  • Session arena (SMH-PLAN-82POBOKB): Entry is a Copy record over one Vec<u8> arena, content is canonical CBOR bytes decoded on demand, frames embed content as one byte string (VERSION_ENTRY = 2): 2002 → 3, load constant, turn 220 → 208.
  • Arc<[ToolMetadata]> registry snapshot on ProviderRequest.tools instead of a per-request sorted clone of every schema: 208 → 42.
  • Provider records over bounded mpsc, single-pass masked request build: 222 → 220 in-process (the first turn per process pays +1 allocation of lazy global init; bounds are steady state after a warm-up).
  • HTTP body channel bounded (16 chunks) and terminated by an explicit HttpBodyEnd; RPC worker owns the agent, reader owns responses, two atomics plus channels.

Decisions taken on the way: bundled Mozilla roots for TLS (platform-verifier deferred); record_channel bound 64 with a documented per-round invariant; smith-alloc counters stay global atomics (≈ 12 ns per allocation) until a hot path shows the cost; per-frame content Vec on load accepted until the reader hands slices.

Remaining floor of the mock turn (42): request encoding, stream event vectors, two entry encodes and frame writes; next candidates are a slice-handing frame reader, a borrowed EntryContent view at record time, and a per-session text pool.

Open questions

  • Kernel pool sizing: lazy pool capped at cores vs one thread per plugin.
  • Render tick: fixed 60 Hz vs on-demand with deadline.
  • Sealed WASI breadth: which further imports (random seed, wall-clock, filesystem) get Smith substitutes before a non-trivial guest std program runs; monotonic_now is still never injected.
  • Per-turn peak: reset or epoch API on smith-alloc; thread-local counters.
  • Trust store: revisit platform-verifier when a corporate proxy case appears.

Sources

  • Code and lockfile at tztonmyo (dd069f92); spike changes qwzlzyqk, lnrvruyv, onllulxs; promoted stack tlszyzws through wqptrptw.
  • Registry sources: wasmtime-wasi 49.0.1, wasmtime-wasi-io 49.0.1, ureq 3.4.2, ureq-proto 0.6.4.
  • SMH-SPEC-WASM0001 (sealed WASI law, single-threaded instances), SMH-SPEC-SPEC0001 (bounded frame allocation).
  • SMH-RESEARCH-RSCH0001 (arena patterns), SMH-RESEARCH-RSCH0010 (session size distribution: p99 ≈ 9 MiB JSONL, compaction near 250k tokens).
---
id: SMH-RESEARCH-4KD6V6IH
type: research
title: "Concurrency and Memory Shape"
---

# Concurrency and Memory Shape

Date: 2026-09-26.
Status: evidence from code survey and three in-repo spikes; candidates, not mandates.
Supersedes the runtime-related parts of SMH-RESEARCH-RSCH0001 and SMH-RESEARCH-RSCH0010 (both predate the wasm plugin law and assume Lua and an async provider boundary).

## Question

Which thread, runtime, ownership, and allocation shape gives Smith a render loop that never stalls, plugins that run as isolated kernels, bounded network and file I/O, deterministic replay, and cache-efficient data, while staying small enough that later web, daemon, and p2p frontends attach without reshaping the core?

## Facts

### Code survey at `tztonmyo` (≈11.5k loc host crates)

| crate | loc | `.clone()` | `Arc<` | `Box<dyn` | lifetimes | `Cow` | `async fn` |
|---|---|---|---|---|---|---|---|
| smith | 1275 | 3 | 4 | 0 | 1 | 0 | 0 |
| smith-core | 5201 | 68 | 4 | 0 | 0 | 0 | 0 |
| smith-ai | 3356 | 15 | 8 | 2 | 0 | 0 | 0 |
| smith-rpc | 487 | 9 | 1 | 0 | 0 | 0 | 0 |
| smith-harness | 678 | 7 | 2 | 0 | 1 | 0 | 0 |
| smith-cli | 451 | 2 | 0 | 0 | 0 | 0 | 0 |

Observed hot spots:

- `smith-core/src/session.rs` `ancestry`: builds a fresh `HashMap<EntryId, &Entry>` per call, then clones every `Entry` up the chain.
  Measured: 2020 allocations, 549 KiB for `fork_at` over 1000 entries (spike C).
- `smith-core/src/agent.rs` request build: clones the full message history per provider turn, then masks in place.
- `Entry.content` embeds `serde_json::Value`: one heap node per JSON node.
- `smith::provider::RecordSpool = Arc<Mutex<Vec<Value>>>`: three indirections plus a lock on the recording path.
- `smith::provider::StreamFn` returns `BoxStream<'static>`; no consumer awaits it, the harness drives it from a thread.
- `smith-harness/src/http.rs` is already blocking: reader thread, `std::sync::mpsc`, `CancelHandle` atomic.
- `clippy::needless_pass_by_value` is allowed workspace-wide.

### Runtime dependency roots at `tztonmyo`

`tokio` enters only through two engines, never by Smith's choice:

```text
tokio 1.53
├── reqwest 0.13 → hyper 1, hyper-util, tower, tokio-util, tokio-rustls, aws-lc-sys
└── wasmtime-wasi 49 (hard dep; feature p2 enables wasmtime/async; sync API is block_on)
```

`wasmtime-wasi` also pulls `rand`, `cap-primitives`, `rustix` (ambient OS surface).
`wasmtime-wasi-io` 49 is `no_std` and tokio-free but hard-enables `wasmtime/async` and still leaves clocks and cli unimplemented.

The echo guest (`wasm32-wasip2`, wit-bindgen 0.62) imports exactly: `wasi:io/{poll,error,streams}`, `wasi:clocks/monotonic-clock`, `wasi:cli/{stdin,stdout,stderr,environment,exit,terminal-*}` at 0.2.9; no filesystem, random, sockets, or wall-clock.
Guest 0.2.9 imports resolve against host 0.2.12 definitions through semver matching.

### Allocator control on stable Rust

| mechanism | stable | effort | fit |
|---|---|---|---|
| `#[global_allocator]` swap or wrapper | yes | one `unsafe impl GlobalAlloc` in a bin or dedicated crate | counting, bounding, attribution of all code including engines |
| `allocator_api` (`Vec::new_in`) | no | nightly; `String` has no `_in`; engines allocate globally anyway | rejected |
| lifetime arena (`bumpalo`) | yes | `'a` infects async, serde, wasm host, plugin API | rejected for core; scratch only where no reference escapes |
| index arena (`Vec<T>` + `u32` handle) | yes | layout discipline only | chosen for session data |

Lifetime grouping by type is almost never the logical grouping; Smith's logical groups are process, session, turn, kernel call, frame, io request.

## Spikes

All three ran in-repo as sibling changes off `tztonmyo` (shared build cache); each is a candidate for promotion after spec coverage.

### A. Sealed WASI shims replace `wasmtime-wasi` (`qwzlzyqk` / `db1b1af4`)

Result: supported.

- One `wasmtime::component::bindgen!` call covers the plugin world plus an inline `smith:host` world importing exactly the guest's 12 WASI interfaces; generates sync `Host` traits; `exit` marked trappable.
- Shims in `smith-harness/src/wasm/sealed_wasi.rs`: empty environment/arguments/cwd, stdin closed, stdout/stderr accept and discard, terminals `None`, exit traps, monotonic clock returns host-injected value (0 at load), every pollable ready, no `ResourceTable`, no `WasiCtx`.
- WIT: `wit/deps/{io,clocks,cli}.wit` copied from wasmtime-wasi 49.0.1 (Apache-2.0 WITH LLVM-exception); `cli.wit` has its `imports`/`command` worlds removed because they reference filesystem, sockets, random; header states the modification.
- Gates: `cargo x check|test|lint -p smith-harness` green, 12 tests incl. echo end to end; negative control (drop one import) fails instantiation as expected.
- Tree: 298 → 277 crates; `wasmtime feature "async"` present → absent; `cap-primitives`, `rand`, `wasmtime-wasi` gone; `tokio` remains only via reqwest.
- Risks: other guests importing random/filesystem/sockets/wall-clock fail at instantiation until further sealed shims land; nothing recorded to the session trace yet; clock frozen; `smith-cli/tests/architecture.rs` still names `wasmtime-wasi`; pre-existing bench failure (`criterion::black_box` deprecated) untouched.

### B. `ureq` 3 HTTP executor (`lnrvruyv` / `057ac857`)

Result: supported.

- `UreqHttpExecutor` behind the existing `HttpExecutor` trait; shared `admit()` and `stream_body()` pump used by both engines; cancel checked after each read, before send.
- `ureq = { version = "3.4", default-features = false, features = ["rustls"] }` → ureq 3.4.2, ureq-proto 0.6.4, ring, webpki-roots.
- Loopback SSE test (3 events, 200 ms apart), both engines: first chunk 1.55 ms (ureq) vs 0.84 ms (reqwest); last 401 ms both; cancel → channel closed ≤ 200 ms; body cap honoured at 64 MiB.
- Gap timeout (300 ms, server stalls 1 s): body ends at 605 ms.
  Built as a 40-line `GapTimeout` connector over `ureq::unversioned::transport`, which is semver-unstable across minor releases.
  ureq has no stable per-read timeout; `timeout_recv_body` is a whole-body budget.
- reqwest's blocking client had no per-chunk gap timeout; its per-request `.timeout` acts on send and each body read.
- Tree with reqwest removed: 298 → 271 crates (−27 raw, −15 unique normal packages, −45 lockfile entries); removes hyper, tower, tokio-util, hyper-util, tokio-rustls, aws-lc-sys.
  `reqwest` is referenced only in `smith-harness/src/http.rs`, `smith-harness/Cargo.toml`, root `Cargo.toml`.
- TLS trust store changes from platform verifier to bundled Mozilla roots unless `platform-verifier` is enabled.
- Redirect, proxy env, and user-agent defaults not compared.

### C. Scope-tagged counting allocator (`onllulxs` / `e124927f`)

Result: supported.

- New `smith-alloc` library crate (unsafe permitted, `// SAFETY:` comments, 5 unit tests): `Scope` enum (Process, Session, Turn, KernelCall, Frame, IoRequest), RAII `enter`, `snapshot`, `#[global_allocator]` set in `smith-cli`.
- Design: `thread_local!` with `const { Cell<[Scope; 16] + len> }` (Copy, no Drop, never allocates, no TLS destructor); per-scope `AtomicU64 × 4` (allocations, bytes, live, peak) Relaxed; `align`-byte header per block carrying the scope tag so `dealloc`/`realloc` charge the allocating scope; guard restores `min(len, previous)`; falls back to Process during TLS teardown.
- No locks, no allocation in the alloc path; reentrancy verified by exact unit counts.
- Mock eval turn, XDG isolated, three runs identical: Turn 222 allocations / 20 552 B / peak 17 359 B; Process 573 / 134 446 B / peak 70 056 B.
  Without XDG isolation counts drift because `eval.smh` grows between runs.
- `fork_at` over 1000 entries: 2020 allocations, 549 438 B, peak 334 946 B; two consecutive forks equal.
- Overhead: whole mock eval process −1 % … +8 % (noise); alloc-heavy microbench +12 % ≈ 12 ns per allocation over ≈ 53 ns; four atomic RMWs per alloc dominate, thread-local counters would cut it.
- Risks: header changes memory layout (allocator comparisons must account for it); parallel tests sharing a scope mix counters; peak is lifetime not per-guard; spike-only `--cfg smith_alloc_off` toggle and `SMITH_ALLOC_STATS` stderr dump must go; `smith-alloc` missing from xtask `CRATES`; RULES unsafe law and spec coverage needed.

## Tradeoffs

- Async core vs sync core: async buys nothing at Smith's concurrency scale (a handful of provider streams, one bash process, sync grep) and costs `Send + 'static` on every seam, `Arc` infection, and scheduler-dependent ordering inside the determinism boundary.
  Sync core with thread roles makes replay a pure function over the recorded inbound log.
- Runtime crate choice is moot while engines dictate it; the actionable law is confinement (runtime never leaks past the engine-owner crate), not prohibition.
  p2p (`iroh`) and a web frontend will bring tokio back inside their own crates.
- `wasmtime-wasi-io` vs own shims: the former keeps `wasmtime/async` on and covers only io; own shims are ≈ 300 lines, fully sync, and match the spec's "Smith implementations" wording, at the cost of maintaining a hand-stripped `cli.wit`.
- `ureq` vs reqwest: −27 crates, no runtime, ring instead of aws-lc; gap timeout depends on an unstable API and trust-store semantics change.
  reqwest keeps hyper's maturity and platform roots at the cost of the whole tokio stack.
- Global counting allocator: ≈ 12 ns per allocation and one header per block; buys attribution of engine allocations to logical scopes and deterministic count tests.
  Bump-routing scope allocations through the global allocator is unsafe (engine caches, `thread_local!`, lazy statics escape scopes); real frees stay, bump-reset applies only to arenas Smith owns.
- Hard ceilings in a global allocator abort on infallible `Vec::push` inside engines; ceilings are therefore soft in production, hard only in tests and where Smith calls `try_reserve`.

## Candidates

### Concurrency shape

```text
core = pure fn (session arena, ordered inbound event log) → (session arena', outbound effects)
```

| thread | owns | never |
|---|---|---|
| input | terminal events → coalesce → ui channel | blocks on core |
| render | terminal, frame budget, drains ui channel | I/O, waits on producers |
| core | Session arena, agent loop, single consumer of inbound channel | races, ambient time or randomness |
| kernel pool | one plugin instance pinned per worker, sync wasmtime, fuel + epoch | blocking host calls inside guest |
| io | net + fs on blocking threads, bounded chunks, cancel by atomic | render-thread work |

Edges: typed bounded channels; payloads owned or `Arc<immutable>`; no `Mutex` in library crates; `'static` boundaries take snapshots, never force upstream cloning.

Kernel = `(instance, call, inputs) → outputs + effect requests`; effects are dispatched by core to io and re-enter as the next kernel call, so every effect is recorded and replayable and one instance is never on two threads.

Render: drain all pending events, then one frame; on-demand with a 16 ms deadline; layout from the wasm layout plugin runs on the kernel pool and returns an owned tree; producers block on full channels, the renderer never does.

Later frontends: web (own HTTP server runtime in its crate), daemon (listener plus one core thread per session), p2p (`iroh`, tokio confined to that crate).

### Lifetime groups

| group | lives | dies | contents |
|---|---|---|---|
| process | start | exit | config, model catalog, compiled wasm modules, tool registry |
| session | open | close | entry arena, text pool, payload bytes |
| turn | user input | provider stop | request buffers, response accumulation, tool-round scratch |
| kernel call | host→guest | guest returns | guest linear memory, host conversion buffers |
| frame | tick | paint done | layout tree, cell buffer, diff |
| io request | send | last chunk | chunk buffers, decoder state |

### Data layout

- Session entries: contiguous `Vec<Entry>` per branch, fixed-size `Entry` with `u32` spans into a per-session text pool and a payload byte arena; `EntryId(Uuid)` stays the at-rest identity, `EntryIdx(u32)` the in-memory handle; ancestry becomes an index chase.
- Payloads (tool input, provider bodies): canonical bytes at rest, parsed on demand into turn scope.
- Frame: double-buffered `Vec<Cell>`, never reallocated after the first frame.
- Turn scratch: one bump `Vec<u8>`; provider request serialized straight into it; `clear()` per turn.
- Underlying allocator (`System` vs `mimalloc`) decided by benchmark once count tests exist; header cost of the counter must be excluded from that comparison.

### Promotion order

1. Spike A (sealed WASI) after spec coverage of sealed semantics and trace recording; fix architecture test naming.
2. Spike C (`smith-alloc`) after RULES unsafe exception and spec coverage; drop spike toggles; add to xtask `CRATES`; switch counters to thread-local.
3. Spike B (`ureq`) after deciding trust-store policy and pinning `~3.4`; delete reqwest engine and workspace entry.
4. Hot-spot refactors with count bounds: `ancestry` → request build → spool → `Value` payloads.

## Promotion results (2026-09-26, stack `tlszyzws..wqptrptw`)

All three spikes promoted after spec text landed (SMH-SPEC-WASM0001 Sealed WASI; SMH-SPEC-SPEC0001 network transport, memory accounting, index arena).

| measure | before | after |
|---|---|---|
| crates in normal graph | 298 | 210 |
| lockfile packages | 370 | 294 |
| async runtime in graph | tokio via reqwest and wasmtime-wasi | none |
| `Mutex` in library production code | 5 sites | 0, enforced by `cargo x arch` |
| `fork_at(1000)` allocations | 2020 | 3 |
| `Session::from_frames(1000)` allocations | unmeasured | 5, equal at 100 entries |
| mock eval turn allocations | 222 | 42 |

What moved each number:

- `ancestry` as branch prefix: 2020 → 2002 (the floor of owned `Entry` clones).
- Session arena (SMH-PLAN-82POBOKB): `Entry` is a `Copy` record over one `Vec<u8>` arena, content is canonical CBOR bytes decoded on demand, frames embed content as one byte string (`VERSION_ENTRY = 2`): 2002 → 3, load constant, turn 220 → 208.
- `Arc<[ToolMetadata]>` registry snapshot on `ProviderRequest.tools` instead of a per-request sorted clone of every schema: 208 → 42.
- Provider records over bounded `mpsc`, single-pass masked request build: 222 → 220 in-process (the first turn per process pays +1 allocation of lazy global init; bounds are steady state after a warm-up).
- HTTP body channel bounded (16 chunks) and terminated by an explicit `HttpBodyEnd`; RPC worker owns the agent, reader owns responses, two atomics plus channels.

Decisions taken on the way: bundled Mozilla roots for TLS (`platform-verifier` deferred); `record_channel` bound 64 with a documented per-round invariant; `smith-alloc` counters stay global atomics (≈ 12 ns per allocation) until a hot path shows the cost; per-frame content `Vec` on load accepted until the reader hands slices.

Remaining floor of the mock turn (42): request encoding, stream event vectors, two entry encodes and frame writes; next candidates are a slice-handing frame reader, a borrowed `EntryContent` view at record time, and a per-session text pool.

## Open questions

- Kernel pool sizing: lazy pool capped at cores vs one thread per plugin.
- Render tick: fixed 60 Hz vs on-demand with deadline.
- Sealed WASI breadth: which further imports (random seed, wall-clock, filesystem) get Smith substitutes before a non-trivial guest std program runs; `monotonic_now` is still never injected.
- Per-turn peak: reset or epoch API on `smith-alloc`; thread-local counters.
- Trust store: revisit `platform-verifier` when a corporate proxy case appears.

## Sources

- Code and lockfile at `tztonmyo` (`dd069f92`); spike changes `qwzlzyqk`, `lnrvruyv`, `onllulxs`; promoted stack `tlszyzws` through `wqptrptw`.
- Registry sources: `wasmtime-wasi 49.0.1`, `wasmtime-wasi-io 49.0.1`, `ureq 3.4.2`, `ureq-proto 0.6.4`.
- SMH-SPEC-WASM0001 (sealed WASI law, single-threaded instances), SMH-SPEC-SPEC0001 (bounded frame allocation).
- SMH-RESEARCH-RSCH0001 (arena patterns), SMH-RESEARCH-RSCH0010 (session size distribution: p99 ≈ 9 MiB JSONL, compaction near 250k tokens).