SoA vs AoS → array of structs (AoS) interleaves fields; struct of arrays (SoA) separates by field. SoA improves SIMD throughput and cache utilization when iterating one field. convert hot loops to SoA.
hot/cold splitting → separate frequently-accessed data (hot) from rarely-accessed (cold). reduces cache pressure; cold data doesn't pollute cache lines.
cache line padding → align data to cache line boundaries (64 bytes typically). prevents false sharing between threads and avoids partial line loads.
loop tiling/blocking → process data in blocks that fit cache rather than streaming entire dataset. improves temporal locality for multi-dimensional access.
software prefetching → explicitly load future data ahead of need. hides memory latency for pointer-chasing or irregular access patterns.
locality is king → sequential access orders of magnitude faster than random
allocations aren't free → CPU + reclamation time, even when allocator is fast
branch mispredictions expensive → predictable control flow faster
Pitfalls
over-engineering for hardware that may change → focus on timeless principles (locality, independence)
don't defeat compiler → write code compiler/backend can optimize
premature opt → mechanical sympathy guides design, not micro-optimization
SoA complexity → only convert when iterating single field; AoS is fine for record-style access
# Mechanical Sympathy
substrate awareness (runtime, VM, CPU, network, API) → structure data + instructions efficiently.
## Core Concepts
- **memory access** → cache hierarchy friendly vs cache-missing. sequential over random, small working sets
- **allocations** → runtime reclamation tax: short-lived values, fragmentation, excessive retention
- **compiler/backend** → what it can/cannot optimize. verify with output, don't guess
- **patterns** → data-oriented design, dependency-chain-free loops, branch-prediction-friendly code
## Data-Oriented Design
- **SoA vs AoS** → array of structs (AoS) interleaves fields; struct of arrays (SoA) separates by field. SoA improves SIMD throughput and cache utilization when iterating one field. convert hot loops to SoA.
- **hot/cold splitting** → separate frequently-accessed data (hot) from rarely-accessed (cold). reduces cache pressure; cold data doesn't pollute cache lines.
- **cache line padding** → align data to cache line boundaries (64 bytes typically). prevents false sharing between threads and avoids partial line loads.
- **loop tiling/blocking** → process data in blocks that fit cache rather than streaming entire dataset. improves temporal locality for multi-dimensional access.
- **software prefetching** → explicitly load future data ahead of need. hides memory latency for pointer-chasing or irregular access patterns.
## Approach
1. understand substrate → cache hierarchy, memory layout, reclamation, compiler/backend pipeline
2. structure data → contiguous over linked, padding for alignment, SoA for hot loops
3. structure instructions → independent for parallel execution
4. verify compiler/backend output → inspect IR/assembly
5. measure → confirm substrate-aware changes help; don't assume
## Timeless Principles
- measure, don't guess → compiler/backend output + profiling > intuition
- locality is king → sequential access orders of magnitude faster than random
- allocations aren't free → CPU + reclamation time, even when allocator is fast
- branch mispredictions expensive → predictable control flow faster
## Pitfalls
- over-engineering for hardware that may change → focus on timeless principles (locality, independence)
- don't defeat compiler → write code compiler/backend can optimize
- premature opt → mechanical sympathy guides design, not micro-optimization
- SoA complexity → only convert when iterating single field; AoS is fine for record-style access