# Benchmark: batch hot path ## Scope and exclusions In-process hot path of the `firefox_bidi` tool: per-frame processing (`runBatch`) and payload serialization, plus its interaction with real Firefox round-trips. Excluded: browser launch and model latency — they dominate end-to-end wall time and are not this repo's code. ## Measurements Two matched benchmarks: 1. **Microbenchmark** (`__bench__/runBatch.bench.mts`): fake socket, fixture batch of 25 frames (22 small + 3 × 512 KB results). Isolates CPU cost (interpolation, serialization, ledger, spool writes) from the network. 2. **E2E benchmark** (`__bench__/e2e.mts`): real headless Firefox, one `/session` socket, real round-trips; 5-frame batches with three real 300 KB `script.evaluate` results that hit the spool path. Both: interleaved A/B pairs (drift cancellation), fixture/launch setup outside the timer, output-equality gates, Wilcoxon signed-rank on paired differences (normal approximation). ## Findings Microbenchmark (before optimization): `JSON.stringify` of a 512 KB response ≈ 0.11 ms; the old loop serialized every response twice (size check + full payload re-stringify) ≈ 0.36 ms/batch; spool `writeFile` calls sat on the frame critical path. n=30 paired: no significant difference (p=0.69) — fs writes dominated the loop. E2E (the decisive measurement): spool writes deferred to a post-loop drain (`Promise.all`) so they overlap subsequent frame round-trips: | run | inline-spool median | deferred-spool median | paired diff | p | |---|---|---|---|---| | 12 pairs | 55.2 ms | 27.3 ms | 26.7 ms | 0.008 | | 15 pairs | 53.7 ms | 25.3 ms | 24.5 ms | 0.0022 | ≈ **46% faster batches** containing oversized payloads. Additional micro optimizations kept: single serialization per response (size check shares the payload text), `interpolate` fast-path for strings without `{{`. Follow-up (profiled, 2026-09-26): V8 `--cpu-prof` over the micro loop showed ~49% of samples in the two `runBatch` implementations' own bodies and only ~5% in callees; a controlled experiment (`big=3` vs `big=0` fixture) pinned ~94% of candidate hot-path cost on the oversized-result path, where each big frame paid **two** 512 KB serializations: the full-frame size probe plus the spool payload. Fix: serialize the result first. `resultText` decides the spool path and is written verbatim when oversized — one big serialization per frame. The old marker semantics (`size` = pre-marker full-frame length) are preserved exactly by wrapping a `0` placeholder result once (small stringify) and substituting its span: `fullLength = probe.length - 1 + resultText.length`. Borderline threshold uses `MAX_INLINE - 1024` wrapper slack so marked frames stay under the inline cap. | run | two-pass median | single-pass median | paired diff | p | |---|---|---|---|---| | micro n=30 | 3.11 ms | 1.89 ms | ≈1.2 ms | 1.7e-6 | | e2e n=12 | 58.4 ms | 27.7 ms | ≈30 ms | 0.0047 | | e2e replicate | 56.2 ms | 30.2 ms | ≈26 ms | 0.015 | | e2e replicate | 56.8 ms | 27.9 ms | ≈29 ms | 0.0096 | ≈ **50% faster batches** with oversized payloads, replicated; output gate (bit-identical payloads vs frozen baseline, including `marker.size`) green. Small results serialize twice (probe + wrapper) — measured free (the `big=0` run medians 0.14 ms for 25 frames). ## Decision Deferred spool drain is behavior-preserving (files are flushed before the tool returns; write errors still surface) and measured significant against a real browser. Serialization and interpolation changes are measured neutral to positive and strictly reduce work. ## Reproduce ```sh mise run //extensions/firefox-bidi:bench mise run //extensions/firefox-bidi:bench-e2e -- 15 ```