In-process hot path of the firefox_bidi tool: per-frame processing
(runBatch) and payload serialization, plus its interaction with real
Firefox round-trips. Excluded: browser launch and model latency — they
dominate end-to-end wall time and are not this repo's code.
Measurements
Two matched benchmarks:
Microbenchmark (__bench__/runBatch.bench.mts): fake socket, fixture
batch of 25 frames (22 small + 3 × 512 KB results). Isolates CPU cost
(interpolation, serialization, ledger, spool writes) from the network.
E2E benchmark (__bench__/e2e.mts): real headless Firefox, one
/session socket, real round-trips; 5-frame batches with three real
300 KB script.evaluate results that hit the spool path.
Both: interleaved A/B pairs (drift cancellation), fixture/launch setup
outside the timer, output-equality gates, Wilcoxon signed-rank on paired
differences (normal approximation).
Findings
Microbenchmark (before optimization): JSON.stringify of a 512 KB response
≈ 0.11 ms; the old loop serialized every response twice (size check + full
payload re-stringify) ≈ 0.36 ms/batch; spool writeFile calls sat on the
frame critical path. n=30 paired: no significant difference (p=0.69) — fs
writes dominated the loop.
E2E (the decisive measurement): spool writes deferred to a post-loop drain
(Promise.all) so they overlap subsequent frame round-trips:
run
inline-spool median
deferred-spool median
paired diff
p
12 pairs
55.2 ms
27.3 ms
26.7 ms
0.008
15 pairs
53.7 ms
25.3 ms
24.5 ms
0.0022
≈ 46% faster batches containing oversized payloads. Additional micro
optimizations kept: single serialization per response (size check shares the
payload text), interpolate fast-path for strings without {{.
Follow-up (profiled, 2026-09-26): V8 --cpu-prof over the micro loop showed
~49% of samples in the two runBatch implementations' own bodies and only
~5% in callees; a controlled experiment (big=3 vs big=0 fixture) pinned
~94% of candidate hot-path cost on the oversized-result path, where each big
frame paid two 512 KB serializations: the full-frame size probe plus the
spool payload.
Fix: serialize the result first. resultText decides the spool path and is
written verbatim when oversized — one big serialization per frame. The old
marker semantics (size = pre-marker full-frame length) are preserved
exactly by wrapping a 0 placeholder result once (small stringify) and
substituting its span: fullLength = probe.length - 1 + resultText.length.
Borderline threshold uses MAX_INLINE - 1024 wrapper slack so marked frames
stay under the inline cap.
run
two-pass median
single-pass median
paired diff
p
micro n=30
3.11 ms
1.89 ms
≈1.2 ms
1.7e-6
e2e n=12
58.4 ms
27.7 ms
≈30 ms
0.0047
e2e replicate
56.2 ms
30.2 ms
≈26 ms
0.015
e2e replicate
56.8 ms
27.9 ms
≈29 ms
0.0096
≈ 50% faster batches with oversized payloads, replicated; output gate
(bit-identical payloads vs frozen baseline, including marker.size) green.
Small results serialize twice (probe + wrapper) — measured free (the
big=0 run medians 0.14 ms for 25 frames).
Decision
Deferred spool drain is behavior-preserving (files are flushed before the
tool returns; write errors still surface) and measured significant against a
real browser. Serialization and interpolation changes are measured neutral
to positive and strictly reduce work.
Reproduce
mise run //extensions/firefox-bidi:bench
mise run //extensions/firefox-bidi:bench-e2e -- 15
# Benchmark: batch hot path
## Scope and exclusions
In-process hot path of the `firefox_bidi` tool: per-frame processing
(`runBatch`) and payload serialization, plus its interaction with real
Firefox round-trips. Excluded: browser launch and model latency — they
dominate end-to-end wall time and are not this repo's code.
## Measurements
Two matched benchmarks:
1. **Microbenchmark** (`__bench__/runBatch.bench.mts`): fake socket, fixture
batch of 25 frames (22 small + 3 × 512 KB results). Isolates CPU cost
(interpolation, serialization, ledger, spool writes) from the network.
2. **E2E benchmark** (`__bench__/e2e.mts`): real headless Firefox, one
`/session` socket, real round-trips; 5-frame batches with three real
300 KB `script.evaluate` results that hit the spool path.
Both: interleaved A/B pairs (drift cancellation), fixture/launch setup
outside the timer, output-equality gates, Wilcoxon signed-rank on paired
differences (normal approximation).
## Findings
Microbenchmark (before optimization): `JSON.stringify` of a 512 KB response
≈ 0.11 ms; the old loop serialized every response twice (size check + full
payload re-stringify) ≈ 0.36 ms/batch; spool `writeFile` calls sat on the
frame critical path. n=30 paired: no significant difference (p=0.69) — fs
writes dominated the loop.
E2E (the decisive measurement): spool writes deferred to a post-loop drain
(`Promise.all`) so they overlap subsequent frame round-trips:
| run | inline-spool median | deferred-spool median | paired diff | p |
|---|---|---|---|---|
| 12 pairs | 55.2 ms | 27.3 ms | 26.7 ms | 0.008 |
| 15 pairs | 53.7 ms | 25.3 ms | 24.5 ms | 0.0022 |
≈ **46% faster batches** containing oversized payloads. Additional micro
optimizations kept: single serialization per response (size check shares the
payload text), `interpolate` fast-path for strings without `{{`.
Follow-up (profiled, 2026-09-26): V8 `--cpu-prof` over the micro loop showed
~49% of samples in the two `runBatch` implementations' own bodies and only
~5% in callees; a controlled experiment (`big=3` vs `big=0` fixture) pinned
~94% of candidate hot-path cost on the oversized-result path, where each big
frame paid **two** 512 KB serializations: the full-frame size probe plus the
spool payload.
Fix: serialize the result first. `resultText` decides the spool path and is
written verbatim when oversized — one big serialization per frame. The old
marker semantics (`size` = pre-marker full-frame length) are preserved
exactly by wrapping a `0` placeholder result once (small stringify) and
substituting its span: `fullLength = probe.length - 1 + resultText.length`.
Borderline threshold uses `MAX_INLINE - 1024` wrapper slack so marked frames
stay under the inline cap.
| run | two-pass median | single-pass median | paired diff | p |
|---|---|---|---|---|
| micro n=30 | 3.11 ms | 1.89 ms | ≈1.2 ms | 1.7e-6 |
| e2e n=12 | 58.4 ms | 27.7 ms | ≈30 ms | 0.0047 |
| e2e replicate | 56.2 ms | 30.2 ms | ≈26 ms | 0.015 |
| e2e replicate | 56.8 ms | 27.9 ms | ≈29 ms | 0.0096 |
≈ **50% faster batches** with oversized payloads, replicated; output gate
(bit-identical payloads vs frozen baseline, including `marker.size`) green.
Small results serialize twice (probe + wrapper) — measured free (the
`big=0` run medians 0.14 ms for 25 frames).
## Decision
Deferred spool drain is behavior-preserving (files are flushed before the
tool returns; write errors still surface) and measured significant against a
real browser. Serialization and interpolation changes are measured neutral
to positive and strictly reduce work.
## Reproduce
```sh
mise run //extensions/firefox-bidi:bench
mise run //extensions/firefox-bidi:bench-e2e -- 15
```