Conversation
Benchmark:
|
| Scenario | Base req/s | Head req/s | Head / base | Rounds | Noise | Busy | CPU per request |
|---|---|---|---|---|---|---|---|
| http/hello-world | 113392 | 112846 | 0.996x (-0.4%) | 0.99 to 1.00 | ±1.4% | 100% / 100% | 8.81 / 8.85 us (1.004x) |
| http/headers | 96150 | 95283 | 0.990x (-1.0%) | 0.98 to 1.00 | ±1.4% | 100% / 100% | 10.39 / 10.49 us (1.011x) |
| http/json-post | 79780 | 84144 | 1.053x (+5.3%) | 1.04 to 1.06 | ±5.7% | 100% / 100% | 12.52 / 11.87 us (0.950x) |
| http/cached | 127653 | 127095 | 0.994x (-0.6%) | 0.99 to 1.00 | ±0.8% | 100% / 100% | 7.83 / 7.86 us (1.006x) |
| ws/echo-20b | 118085 | 117008 | 0.991x (-0.9%) | 0.98 to 1.00 | ±2.0% | 100% / 100% | 8.46 / 8.54 us (1.009x) |
| ws/echo-4kb | 100852 | 102354 | 1.017x (+1.7%) | 1.00 to 1.02 | ±1.9% | 100% / 100% | 9.90 / 9.76 us (0.983x) |
4 rounds of 3s per scenario, each round loading base, head and a second process of each one, from a copy of its build, after the other, in an order that changes every round. "Head / base" is the median of the per-round ratios of req/s and "Rounds" their range. "Noise" is how far base/base and head/head, the same code on both sides, got from 1 in this same run: that is what the machine did, so a row is marked only when the median moved further than that, at least 2%, and every round moved the same way: 👀 slower, 🏆 faster. "Busy" is the cpu time of the server process over the round and "CPU per request" that time divided by the requests it answered, with the head/base ratio judged against its own same-code band and bold when it cleared it. When busy is well under 100% the load generator set the pace, the req/s are its and the cpu per request is the column to read. Only the ratios are comparable across runs, the absolute figures depend on the runner.
Node v26.10.0, AMD EPYC 7763 64-Core Processor, 4 cores (server on cpu 0, load on 2,3), http_load_test and load_test with 100 connections.
Only the binary path optimization from #1302.
Since Node.js 20 the
ArrayBuffer::Data()function can replaceArrayBuffer::GetBackingStore()->Data().It removes the need to instantiate the
shared_ptrand increase performance.This optimization is possible because the original code does not keep the
shared_ptrand use the data synchronously.Test case: 10 x
writeHeader(32 bytes, 32 bytes)152k req/sec -> 155k req/sec -> +2%