Scope: in-process handling of 16 labeled CDP replies of 512 KiB each with a fake transport, including JSON serialization and private spool writes.
Browser launch, network latency, capture encoding, and small replies are excluded.
Run mise run //extensions/chrome-cdp:bench -- 60 compare for interleaved old-cost/current pairs, alternating order each pair.
Run mise run //extensions/chrome-cdp:profile for a Node V8 CPU profile under .pi/tmp/chrome-cdp-bench/profiles.
The script creates and removes a fresh spool directory for each sample; five untimed warmups precede sampling.
It checks every frame is spooled; __tests__/events.test.ts checks exact file bytes, metadata, and small-response label caching.
The compare mode reproduces the former extra label-size serialization immediately before returning the fake reply; the current branch skips that extra serialization.
This is an approximation of the former production path, not a second copy of the old implementation.
Each reported sample divides a 16-reply batch by 16, in milliseconds per reply.
The two-sided exact sign test counts per-pair wins; it does not assume normally distributed timings.
Before the change, two independent unpaired full-batch runs gave medians of 0.441 and 0.413 ms; candidate runs gave 0.473 and 0.457 ms.
Those unpaired measurements are inconclusive because filesystem and machine drift exceed the difference.
The interleaved comparison isolates the removed serialization and supports a CPU improvement for large labeled replies, not an end-to-end speed claim.
An unlabeled-spool control had before medians of 0.353 and 0.377 ms and after medians of 0.390 and 0.338 ms; no consistent regression detected.
V8 --cpu-prof before the change sampled executeOperations 320 times and spoolFrame 159 times during the batch workload; both serialized the same labeled response.
The optimization passes the serialized frame from the cache-size check to spoolFrame, preserving exact spool bytes and cache threshold while avoiding the second stringify.
Profile sample counts are diagnostic, not a cross-run performance statistic.
Deferred spooling was not adopted.
It cannot eliminate writes because returned paths must exist before the tool result returns.
It would retain up to the full batch of large replies in memory before checking disk capacity, weakening the 16 MiB quota and making write failure arrive after more work.
Current measurements target the only removed cost with equivalent path, bytes, and quota behavior: redundant serialization before immediate bounded spool writes.
# Chrome CDP batch benchmark
Scope: in-process handling of 16 labeled CDP replies of 512 KiB each with a fake transport, including JSON serialization and private spool writes.
Browser launch, network latency, capture encoding, and small replies are excluded.
Run `mise run //extensions/chrome-cdp:bench -- 60 compare` for interleaved old-cost/current pairs, alternating order each pair.
Run `mise run //extensions/chrome-cdp:profile` for a Node V8 CPU profile under `.pi/tmp/chrome-cdp-bench/profiles`.
The script creates and removes a fresh spool directory for each sample; five untimed warmups precede sampling.
It checks every frame is spooled; `__tests__/events.test.ts` checks exact file bytes, metadata, and small-response label caching.
The `compare` mode reproduces the former extra label-size serialization immediately before returning the fake reply; the current branch skips that extra serialization.
This is an approximation of the former production path, not a second copy of the old implementation.
Each reported sample divides a 16-reply batch by 16, in milliseconds per reply.
The two-sided exact sign test counts per-pair wins; it does not assume normally distributed timings.
Node v26.9.0, x64 Linux, Ryzen 5 5600X, 60 pairs per run:
| Run | Old-cost median | Current median | Per-run SD old/current | Wins | Sign-test p |
| --- | ---: | ---: | ---: | ---: | ---: |
| 1 | 0.478 ms | 0.289 ms | 0.093/0.065 ms | 48/60 | 3.18e-6 |
| 2 | 0.420 ms | 0.302 ms | 0.133/0.089 ms | 50/60 | 1.62e-7 |
Before the change, two independent unpaired full-batch runs gave medians of 0.441 and 0.413 ms; candidate runs gave 0.473 and 0.457 ms.
Those unpaired measurements are inconclusive because filesystem and machine drift exceed the difference.
The interleaved comparison isolates the removed serialization and supports a CPU improvement for large labeled replies, not an end-to-end speed claim.
An unlabeled-spool control had before medians of 0.353 and 0.377 ms and after medians of 0.390 and 0.338 ms; no consistent regression detected.
V8 `--cpu-prof` before the change sampled `executeOperations` 320 times and `spoolFrame` 159 times during the batch workload; both serialized the same labeled response.
The optimization passes the serialized frame from the cache-size check to `spoolFrame`, preserving exact spool bytes and cache threshold while avoiding the second stringify.
Profile sample counts are diagnostic, not a cross-run performance statistic.
Deferred spooling was not adopted.
It cannot eliminate writes because returned paths must exist before the tool result returns.
It would retain up to the full batch of large replies in memory before checking disk capacity, weakening the 16 MiB quota and making write failure arrive after more work.
Current measurements target the only removed cost with equivalent path, bytes, and quota behavior: redundant serialization before immediate bounded spool writes.