Luigit
repositories / pi-ext

pi-ext

bugabingas pi extensions

owned by admin

extensions/firefox-bidi/__bench__/README.md

Raw
Rendered preview

Benchmark: batch hot path

Scope and exclusions

In-process hot path of the firefox_bidi tool: per-frame processing (runBatch) and payload serialization, plus its interaction with real Firefox round-trips. Excluded: browser launch and model latency — they dominate end-to-end wall time and are not this repo's code.

Measurements

Two matched benchmarks:

  1. Microbenchmark (__bench__/runBatch.bench.mts): fake socket, fixture batch of 25 frames (22 small + 3 × 512 KB results). Isolates CPU cost (interpolation, serialization, ledger, spool writes) from the network.
  2. E2E benchmark (__bench__/e2e.mts): real headless Firefox, one /session socket, real round-trips; 5-frame batches with three real 300 KB script.evaluate results that hit the spool path.

Both: interleaved A/B pairs (drift cancellation), fixture/launch setup outside the timer, output-equality gates, Wilcoxon signed-rank on paired differences (normal approximation).

Findings

Microbenchmark (before optimization): JSON.stringify of a 512 KB response ≈ 0.11 ms; the old loop serialized every response twice (size check + full payload re-stringify) ≈ 0.36 ms/batch; spool writeFile calls sat on the frame critical path. n=30 paired: no significant difference (p=0.69) — fs writes dominated the loop.

E2E (the decisive measurement): spool writes deferred to a post-loop drain (Promise.all) so they overlap subsequent frame round-trips:

run inline-spool median deferred-spool median paired diff p
12 pairs 55.2 ms 27.3 ms 26.7 ms 0.008
15 pairs 53.7 ms 25.3 ms 24.5 ms 0.0022

≈ 46% faster batches containing oversized payloads. Additional micro optimizations kept: single serialization per response (size check shares the payload text), interpolate fast-path for strings without {{.

Follow-up (profiled, 2026-09-26): V8 --cpu-prof over the micro loop showed ~49% of samples in the two runBatch implementations' own bodies and only ~5% in callees; a controlled experiment (big=3 vs big=0 fixture) pinned ~94% of candidate hot-path cost on the oversized-result path, where each big frame paid two 512 KB serializations: the full-frame size probe plus the spool payload.

Fix: serialize the result first. resultText decides the spool path and is written verbatim when oversized — one big serialization per frame. The old marker semantics (size = pre-marker full-frame length) are preserved exactly by wrapping a 0 placeholder result once (small stringify) and substituting its span: fullLength = probe.length - 1 + resultText.length. Borderline threshold uses MAX_INLINE - 1024 wrapper slack so marked frames stay under the inline cap.

run two-pass median single-pass median paired diff p
micro n=30 3.11 ms 1.89 ms ≈1.2 ms 1.7e-6
e2e n=12 58.4 ms 27.7 ms ≈30 ms 0.0047
e2e replicate 56.2 ms 30.2 ms ≈26 ms 0.015
e2e replicate 56.8 ms 27.9 ms ≈29 ms 0.0096

≈ 50% faster batches with oversized payloads, replicated; output gate (bit-identical payloads vs frozen baseline, including marker.size) green. Small results serialize twice (probe + wrapper) — measured free (the big=0 run medians 0.14 ms for 25 frames).

Decision

Deferred spool drain is behavior-preserving (files are flushed before the tool returns; write errors still surface) and measured significant against a real browser. Serialization and interpolation changes are measured neutral to positive and strictly reduce work.

Reproduce

mise run //extensions/firefox-bidi:bench
mise run //extensions/firefox-bidi:bench-e2e -- 15
# Benchmark: batch hot path

## Scope and exclusions

In-process hot path of the `firefox_bidi` tool: per-frame processing
(`runBatch`) and payload serialization, plus its interaction with real
Firefox round-trips. Excluded: browser launch and model latency — they
dominate end-to-end wall time and are not this repo's code.

## Measurements

Two matched benchmarks:

1. **Microbenchmark** (`__bench__/runBatch.bench.mts`): fake socket, fixture
   batch of 25 frames (22 small + 3 × 512 KB results). Isolates CPU cost
   (interpolation, serialization, ledger, spool writes) from the network.
2. **E2E benchmark** (`__bench__/e2e.mts`): real headless Firefox, one
   `/session` socket, real round-trips; 5-frame batches with three real
   300 KB `script.evaluate` results that hit the spool path.

Both: interleaved A/B pairs (drift cancellation), fixture/launch setup
outside the timer, output-equality gates, Wilcoxon signed-rank on paired
differences (normal approximation).

## Findings

Microbenchmark (before optimization): `JSON.stringify` of a 512 KB response
≈ 0.11 ms; the old loop serialized every response twice (size check + full
payload re-stringify) ≈ 0.36 ms/batch; spool `writeFile` calls sat on the
frame critical path. n=30 paired: no significant difference (p=0.69) — fs
writes dominated the loop.

E2E (the decisive measurement): spool writes deferred to a post-loop drain
(`Promise.all`) so they overlap subsequent frame round-trips:

| run | inline-spool median | deferred-spool median | paired diff | p |
|---|---|---|---|---|
| 12 pairs | 55.2 ms | 27.3 ms | 26.7 ms | 0.008 |
| 15 pairs | 53.7 ms | 25.3 ms | 24.5 ms | 0.0022 |

≈ **46% faster batches** containing oversized payloads. Additional micro
optimizations kept: single serialization per response (size check shares the
payload text), `interpolate` fast-path for strings without `{{`.

Follow-up (profiled, 2026-09-26): V8 `--cpu-prof` over the micro loop showed
~49% of samples in the two `runBatch` implementations' own bodies and only
~5% in callees; a controlled experiment (`big=3` vs `big=0` fixture) pinned
~94% of candidate hot-path cost on the oversized-result path, where each big
frame paid **two** 512 KB serializations: the full-frame size probe plus the
spool payload.

Fix: serialize the result first. `resultText` decides the spool path and is
written verbatim when oversized — one big serialization per frame. The old
marker semantics (`size` = pre-marker full-frame length) are preserved
exactly by wrapping a `0` placeholder result once (small stringify) and
substituting its span: `fullLength = probe.length - 1 + resultText.length`.
Borderline threshold uses `MAX_INLINE - 1024` wrapper slack so marked frames
stay under the inline cap.

| run | two-pass median | single-pass median | paired diff | p |
|---|---|---|---|---|
| micro n=30 | 3.11 ms | 1.89 ms | ≈1.2 ms | 1.7e-6 |
| e2e n=12 | 58.4 ms | 27.7 ms | ≈30 ms | 0.0047 |
| e2e replicate | 56.2 ms | 30.2 ms | ≈26 ms | 0.015 |
| e2e replicate | 56.8 ms | 27.9 ms | ≈29 ms | 0.0096 |

≈ **50% faster batches** with oversized payloads, replicated; output gate
(bit-identical payloads vs frozen baseline, including `marker.size`) green.
Small results serialize twice (probe + wrapper) — measured free (the
`big=0` run medians 0.14 ms for 25 frames).

## Decision

Deferred spool drain is behavior-preserving (files are flushed before the
tool returns; write errors still surface) and measured significant against a
real browser. Serialization and interpolation changes are measured neutral
to positive and strictly reduce work.

## Reproduce

```sh
mise run //extensions/firefox-bidi:bench
mise run //extensions/firefox-bidi:bench-e2e -- 15
```