Skip to content

Optimize NativeString binary path - #1344

Open
ValoChet wants to merge 1 commit into
uNetworking:masterfrom
ValoChet:nativeStringArrayBuffer
Open

ValoChet wants to merge 1 commit into
uNetworking:masterfrom
ValoChet:nativeStringArrayBuffer

Conversation

@ValoChet

@ValoChet ValoChet commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

Only the binary path optimization from #1302.

Since Node.js 20 the ArrayBuffer::Data() function can replace ArrayBuffer::GetBackingStore()->Data().
It removes the need to instantiate the shared_ptr and increase performance.

This optimization is possible because the original code does not keep the shared_ptr and use the data synchronously.

Test case: 10 x writeHeader(32 bytes, 32 bytes)
152k req/sec -> 155k req/sec -> +2%

@github-actions

github-actions Bot commented Oct 1, 2026

Copy link
Copy Markdown

Benchmark: 642ddbf against 32a27de

Scenario Base req/s Head req/s Head / base Rounds Noise Busy CPU per request
http/hello-world 113392 112846 0.996x (-0.4%) 0.99 to 1.00 ±1.4% 100% / 100% 8.81 / 8.85 us (1.004x)
http/headers 96150 95283 0.990x (-1.0%) 0.98 to 1.00 ±1.4% 100% / 100% 10.39 / 10.49 us (1.011x)
http/json-post 79780 84144 1.053x (+5.3%) 1.04 to 1.06 ±5.7% 100% / 100% 12.52 / 11.87 us (0.950x)
http/cached 127653 127095 0.994x (-0.6%) 0.99 to 1.00 ±0.8% 100% / 100% 7.83 / 7.86 us (1.006x)
ws/echo-20b 118085 117008 0.991x (-0.9%) 0.98 to 1.00 ±2.0% 100% / 100% 8.46 / 8.54 us (1.009x)
ws/echo-4kb 100852 102354 1.017x (+1.7%) 1.00 to 1.02 ±1.9% 100% / 100% 9.90 / 9.76 us (0.983x)

4 rounds of 3s per scenario, each round loading base, head and a second process of each one, from a copy of its build, after the other, in an order that changes every round. "Head / base" is the median of the per-round ratios of req/s and "Rounds" their range. "Noise" is how far base/base and head/head, the same code on both sides, got from 1 in this same run: that is what the machine did, so a row is marked only when the median moved further than that, at least 2%, and every round moved the same way: 👀 slower, 🏆 faster. "Busy" is the cpu time of the server process over the round and "CPU per request" that time divided by the requests it answered, with the head/base ratio judged against its own same-code band and bold when it cleared it. When busy is well under 100% the load generator set the pace, the req/s are its and the cpu per request is the column to read. Only the ratios are comparable across runs, the absolute figures depend on the runner.

Node v26.10.0, AMD EPYC 7763 64-Core Processor, 4 cores (server on cpu 0, load on 2,3), http_load_test and load_test with 100 connections.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant