Luigit
repositories / dotfiles

dotfiles

bugabingas dorkfiles

owned by admin

pi/agent/skills/optimize/references/mechanical-sympathy.md

Raw
Rendered preview

Mechanical Sympathy

substrate awareness (runtime, VM, CPU, network, API) → structure data + instructions efficiently.

Core Concepts

  • memory access → cache hierarchy friendly vs cache-missing. sequential over random, small working sets
  • allocations → runtime reclamation tax: short-lived values, fragmentation, excessive retention
  • compiler/backend → what it can/cannot optimize. verify with output, don't guess
  • patterns → data-oriented design, dependency-chain-free loops, branch-prediction-friendly code

Data-Oriented Design

  • SoA vs AoS → array of structs (AoS) interleaves fields; struct of arrays (SoA) separates by field. SoA improves SIMD throughput and cache utilization when iterating one field. convert hot loops to SoA.
  • hot/cold splitting → separate frequently-accessed data (hot) from rarely-accessed (cold). reduces cache pressure; cold data doesn't pollute cache lines.
  • cache line padding → align data to cache line boundaries (64 bytes typically). prevents false sharing between threads and avoids partial line loads.
  • loop tiling/blocking → process data in blocks that fit cache rather than streaming entire dataset. improves temporal locality for multi-dimensional access.
  • software prefetching → explicitly load future data ahead of need. hides memory latency for pointer-chasing or irregular access patterns.

Approach

  1. understand substrate → cache hierarchy, memory layout, reclamation, compiler/backend pipeline
  2. structure data → contiguous over linked, padding for alignment, SoA for hot loops
  3. structure instructions → independent for parallel execution
  4. verify compiler/backend output → inspect IR/assembly
  5. measure → confirm substrate-aware changes help; don't assume

Timeless Principles

  • measure, don't guess → compiler/backend output + profiling > intuition
  • locality is king → sequential access orders of magnitude faster than random
  • allocations aren't free → CPU + reclamation time, even when allocator is fast
  • branch mispredictions expensive → predictable control flow faster

Pitfalls

  • over-engineering for hardware that may change → focus on timeless principles (locality, independence)
  • don't defeat compiler → write code compiler/backend can optimize
  • premature opt → mechanical sympathy guides design, not micro-optimization
  • SoA complexity → only convert when iterating single field; AoS is fine for record-style access
# Mechanical Sympathy

substrate awareness (runtime, VM, CPU, network, API) → structure data + instructions efficiently.

## Core Concepts

- **memory access** → cache hierarchy friendly vs cache-missing. sequential over random, small working sets
- **allocations** → runtime reclamation tax: short-lived values, fragmentation, excessive retention
- **compiler/backend** → what it can/cannot optimize. verify with output, don't guess
- **patterns** → data-oriented design, dependency-chain-free loops, branch-prediction-friendly code

## Data-Oriented Design

- **SoA vs AoS** → array of structs (AoS) interleaves fields; struct of arrays (SoA) separates by field. SoA improves SIMD throughput and cache utilization when iterating one field. convert hot loops to SoA.
- **hot/cold splitting** → separate frequently-accessed data (hot) from rarely-accessed (cold). reduces cache pressure; cold data doesn't pollute cache lines.
- **cache line padding** → align data to cache line boundaries (64 bytes typically). prevents false sharing between threads and avoids partial line loads.
- **loop tiling/blocking** → process data in blocks that fit cache rather than streaming entire dataset. improves temporal locality for multi-dimensional access.
- **software prefetching** → explicitly load future data ahead of need. hides memory latency for pointer-chasing or irregular access patterns.

## Approach

1. understand substrate → cache hierarchy, memory layout, reclamation, compiler/backend pipeline
2. structure data → contiguous over linked, padding for alignment, SoA for hot loops
3. structure instructions → independent for parallel execution
4. verify compiler/backend output → inspect IR/assembly
5. measure → confirm substrate-aware changes help; don't assume

## Timeless Principles

- measure, don't guess → compiler/backend output + profiling > intuition
- locality is king → sequential access orders of magnitude faster than random
- allocations aren't free → CPU + reclamation time, even when allocator is fast
- branch mispredictions expensive → predictable control flow faster

## Pitfalls

- over-engineering for hardware that may change → focus on timeless principles (locality, independence)
- don't defeat compiler → write code compiler/backend can optimize
- premature opt → mechanical sympathy guides design, not micro-optimization
- SoA complexity → only convert when iterating single field; AoS is fine for record-style access