--- name: data-oriented-rust description: "Use whenever a Rust struct, enum, collection, arena, or hot loop's data shape is written or edited: choose layout for cache lines and allocation count, then prove it with size and allocation tests." --- # Data-oriented Rust Memory is the cost; the CPU is free. Shape data for how it is accessed, not for how it is named. ## Before shaping 1. Name the access pattern: which fields, in which order, how many instances, how often. 2. Compute size and alignment by hand, then check with `size_of` (or `-Zprint-type-sizes` on nightly). Alignment is the largest field's; size rounds up to it; every `bool` in a padded struct costs a word. 3. Say which lifetime group owns the data (process, session, turn, frame, request). ## Shrink the hot record - Indexes (`u32` newtype) instead of pointers and `Box`; halves size, drops alignment to 4, no lifetime infection. - Smallest integer that fits; `usize` only at the use site. - Booleans and rare fields out of band: partition into two arrays, or a side table keyed by index for sparse data. - Box the one oversized enum variant; enum size is the largest variant plus tag. - Frozen `Vec` → `Box<[T]>`; often-empty `Vec` inside a hot type → `ThinVec`. - `Arc` | `Arc<[T]>`, never `Arc` | `Arc>`. - Niches are free: `Option`, `Option<&T>`, `Option>` add no bytes. - Types over 128 bytes are copied by `memcpy`; keep hot records under one or two cache lines. ## Choose the layout - Contiguous `Vec` beats any pointer structure for iteration; linked lists and `Vec>` lose an order of magnitude. - Record-at-a-time access → array of structs. Field-at-a-time hot loops → struct of arrays (one `Vec` per field; `soa-rs` when a derive earns it). - Tag checked inside a hot loop → split by variant into separate arrays; the tag becomes which array. - Nested `Vec>` → one flat `Vec` plus offsets. - Generic `T` over `dyn Trait` when the loop body vectorizes; `dyn` costs the indirection, not the vtable. - Sort or group for locality and branch prediction; after sorting owned pointers, rebuild into fresh contiguous storage or an arena so the heap order matches. ## Allocate by lifetime - One arena per lifetime group; reset wholesale; handles are indexes into it. - `with_capacity` from a measured distribution; reuse workhorse collections with `clear()`. - `Cow` only at I/O edges; borrow in flight, own at rest. - Sharing order: move → `&T` → `Arc` → atomic → channel. - Per-thread counters that share a cache line false-share; pad with `#[repr(align(64))]` only there. ## Prove it - Size test: `assert_eq!(size_of::(), N)` under `#[cfg(target_arch = "x86_64")]`; bounds only tighten. - Allocation-count test on the hot path; bounds only tighten. - `criterion` plus `perf stat -d` (cache misses, branch misses, instructions per cycle) before and after; a layout win with more instructions can lose. - Bit-packing and clever encodings only after a measured cache-miss problem; unpacking in the loop has cost the win before. ## Do not - `#[repr(C)]` without FFI or a measured false-sharing case; it forbids the compiler's field reordering. - Shrink cold data; only the most-instantiated types pay back. - Trade a readable owned tree for a packed stream without the benchmark that demands it. Facts, numbers, and sources: [references/techniques.md](references/techniques.md).