Luigit
repositories / smith

smith

There are many coding harnesses - but this one is fast

owned by admin

.pi/skills/data-oriented-rust/SKILL.md

Raw
Rendered preview

name: data-oriented-rust description: "Use whenever a Rust struct, enum, collection, arena, or hot loop's data shape is written or edited: choose layout for cache lines and allocation count, then prove it with size and allocation tests."

Data-oriented Rust

Memory is the cost; the CPU is free. Shape data for how it is accessed, not for how it is named.

Before shaping

  1. Name the access pattern: which fields, in which order, how many instances, how often.
  2. Compute size and alignment by hand, then check with size_of (or -Zprint-type-sizes on nightly). Alignment is the largest field's; size rounds up to it; every bool in a padded struct costs a word.
  3. Say which lifetime group owns the data (process, session, turn, frame, request).

Shrink the hot record

  • Indexes (u32 newtype) instead of pointers and Box; halves size, drops alignment to 4, no lifetime infection.
  • Smallest integer that fits; usize only at the use site.
  • Booleans and rare fields out of band: partition into two arrays, or a side table keyed by index for sparse data.
  • Box the one oversized enum variant; enum size is the largest variant plus tag.
  • Frozen Vec → Box<[T]>; often-empty Vec inside a hot type → ThinVec.
  • Arc<str> | Arc<[T]>, never Arc<String> | Arc<Vec<T>>.
  • Niches are free: Option<NonZeroU32>, Option<&T>, Option<Box<T>> add no bytes.
  • Types over 128 bytes are copied by memcpy; keep hot records under one or two cache lines.

Choose the layout

  • Contiguous Vec<T> beats any pointer structure for iteration; linked lists and Vec<Box<T>> lose an order of magnitude.
  • Record-at-a-time access → array of structs. Field-at-a-time hot loops → struct of arrays (one Vec per field; soa-rs when a derive earns it).
  • Tag checked inside a hot loop → split by variant into separate arrays; the tag becomes which array.
  • Nested Vec<Vec<T>> → one flat Vec<T> plus offsets.
  • Generic T over dyn Trait when the loop body vectorizes; dyn costs the indirection, not the vtable.
  • Sort or group for locality and branch prediction; after sorting owned pointers, rebuild into fresh contiguous storage or an arena so the heap order matches.

Allocate by lifetime

  • One arena per lifetime group; reset wholesale; handles are indexes into it.
  • with_capacity from a measured distribution; reuse workhorse collections with clear().
  • Cow only at I/O edges; borrow in flight, own at rest.
  • Sharing order: move → &T → Arc<immutable> → atomic → channel.
  • Per-thread counters that share a cache line false-share; pad with #[repr(align(64))] only there.

Prove it

  • Size test: assert_eq!(size_of::<Hot>(), N) under #[cfg(target_arch = "x86_64")]; bounds only tighten.
  • Allocation-count test on the hot path; bounds only tighten.
  • criterion plus perf stat -d (cache misses, branch misses, instructions per cycle) before and after; a layout win with more instructions can lose.
  • Bit-packing and clever encodings only after a measured cache-miss problem; unpacking in the loop has cost the win before.

Do not

  • #[repr(C)] without FFI or a measured false-sharing case; it forbids the compiler's field reordering.
  • Shrink cold data; only the most-instantiated types pay back.
  • Trade a readable owned tree for a packed stream without the benchmark that demands it.

Facts, numbers, and sources: references/techniques.md.

---
name: data-oriented-rust
description: "Use whenever a Rust struct, enum, collection, arena, or hot loop's data shape is written or edited: choose layout for cache lines and allocation count, then prove it with size and allocation tests."
---

# Data-oriented Rust

Memory is the cost; the CPU is free.
Shape data for how it is accessed, not for how it is named.

## Before shaping

1. Name the access pattern: which fields, in which order, how many instances, how often.
2. Compute size and alignment by hand, then check with `size_of` (or `-Zprint-type-sizes` on nightly).
   Alignment is the largest field's; size rounds up to it; every `bool` in a padded struct costs a word.
3. Say which lifetime group owns the data (process, session, turn, frame, request).

## Shrink the hot record

- Indexes (`u32` newtype) instead of pointers and `Box`; halves size, drops alignment to 4, no lifetime infection.
- Smallest integer that fits; `usize` only at the use site.
- Booleans and rare fields out of band: partition into two arrays, or a side table keyed by index for sparse data.
- Box the one oversized enum variant; enum size is the largest variant plus tag.
- Frozen `Vec` → `Box<[T]>`; often-empty `Vec` inside a hot type → `ThinVec`.
- `Arc<str>` | `Arc<[T]>`, never `Arc<String>` | `Arc<Vec<T>>`.
- Niches are free: `Option<NonZeroU32>`, `Option<&T>`, `Option<Box<T>>` add no bytes.
- Types over 128 bytes are copied by `memcpy`; keep hot records under one or two cache lines.

## Choose the layout

- Contiguous `Vec<T>` beats any pointer structure for iteration; linked lists and `Vec<Box<T>>` lose an order of magnitude.
- Record-at-a-time access → array of structs.
  Field-at-a-time hot loops → struct of arrays (one `Vec` per field; `soa-rs` when a derive earns it).
- Tag checked inside a hot loop → split by variant into separate arrays; the tag becomes which array.
- Nested `Vec<Vec<T>>` → one flat `Vec<T>` plus offsets.
- Generic `T` over `dyn Trait` when the loop body vectorizes; `dyn` costs the indirection, not the vtable.
- Sort or group for locality and branch prediction; after sorting owned pointers, rebuild into fresh contiguous storage or an arena so the heap order matches.

## Allocate by lifetime

- One arena per lifetime group; reset wholesale; handles are indexes into it.
- `with_capacity` from a measured distribution; reuse workhorse collections with `clear()`.
- `Cow` only at I/O edges; borrow in flight, own at rest.
- Sharing order: move → `&T` → `Arc<immutable>` → atomic → channel.
- Per-thread counters that share a cache line false-share; pad with `#[repr(align(64))]` only there.

## Prove it

- Size test: `assert_eq!(size_of::<Hot>(), N)` under `#[cfg(target_arch = "x86_64")]`; bounds only tighten.
- Allocation-count test on the hot path; bounds only tighten.
- `criterion` plus `perf stat -d` (cache misses, branch misses, instructions per cycle) before and after; a layout win with more instructions can lose.
- Bit-packing and clever encodings only after a measured cache-miss problem; unpacking in the loop has cost the win before.

## Do not

- `#[repr(C)]` without FFI or a measured false-sharing case; it forbids the compiler's field reordering.
- Shrink cold data; only the most-instantiated types pay back.
- Trade a readable owned tree for a packed stream without the benchmark that demands it.

Facts, numbers, and sources: [references/techniques.md](references/techniques.md).