Skip to content

extmod/modlwip: Return a partial or ENOBUFS from a non-blocking send - #19705

Draft
srgg wants to merge 1 commit into
micropython:masterfrom
srgg:modlwip-nonblocking-partial
Draft

srgg wants to merge 1 commit into
micropython:masterfrom
srgg:modlwip-nonblocking-partial

Conversation

@srgg

@srgg srgg commented Sep 16, 2026 •

Copy link
Copy Markdown

Summary

A non-blocking socket.write() blocks for up to 10 s on ERR_MEM (analysis in #19704).

This change makes that path write the prefix that fits — halving write_len and retrying — and return ENOBUFS only when nothing fits, instead of blocking. EAGAIN is not used: POLLOUT reads tcp_sndbuf, so it would busy-spin a select caller; ENOBUFS is a resource error, not would-block. Blocking sockets are unchanged.

Resolves #19704.

Testing

Measured on an OpenMV RT1062 (cyw43 Wi-Fi) streaming MJPEG to a reader throttled to 20 KB/s, on a v1.28.0-based tree: worst write() 1553086 us, 54 events over 100 ms per 320 s; the early return removes both.

Not built against master and run on no port here — the added block copies the adjacent tcp_sndbuf == 0 return, and CI compiles the MICROPY_PY_LWIP ports.

Trade-offs and Alternatives

Halving costs up to log2(tcp_sndbuf) tcp_write attempts on the exhausted pass; returning ENOBUFS with zero bytes is O(1) but makes no progress.

Generative AI

I used generative AI tools when creating this PR, but a human has checked the code and is responsible for the code and the description above.

lwip_tcp_send sizes a write from the tcp_sndbuf counter, but tcp_write
copies into PBUF_RAM from the shared MEM_SIZE heap and returns ERR_MEM
when that heap is exhausted while the counter still shows room. The retry
loop then polls and retries for up to 10s, and its own comment notes that
a non-blocking socket blocks there.

A plain EAGAIN would be wrong: POLLOUT reads tcp_sndbuf, so a select-driven
caller sees the socket writable and busy-spins. On ERR_MEM a non-blocking
socket now shrinks write_len and retries to write the prefix that fits,
returning that partial count, and returns ENOBUFS only when not even one
byte fits -- the resource error, not would-block. Blocking sockets keep the
retry loop.

Signed-off-by: srgg <srggal@gmail.com>
@srgg srgg changed the title fix(modlwip): return a partial or ENOBUFS from a non-blocking send extmod/modlwip: Return a partial or ENOBUFS from a non-blocking send Sep 16, 2026
@codecov

codecov Bot commented Sep 16, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 98.59%. Comparing base (cc12057) to head (9cdfb4e).
⚠️ Report is 4 commits behind head on master.

Additional details and impacted files
@@            Coverage Diff             @@
##           master   #19705      +/-   ##
==========================================
+ Coverage   98.55%   98.59%   +0.03%     
==========================================
  Files         182      182              
  Lines       23335    23335              
  Branches        5        5              
==========================================
+ Hits        22998    23006       +8     
+ Misses        336      328       -8     
  Partials        1        1              
Flag Coverage Δ
unix-coverage-32bit 98.59% <ø> (+0.03%) ⬆️
unix-coverage-64bit 98.52% <ø> (ø)

Flags with carried forward coverage won't be shown. Click here to find out more.

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@github-actions

Copy link
Copy Markdown

Code size report:

Reference:  rp2: Keep machine.RTC ticking while in lightsleep(). [cc12057]
Comparison: extmod/modlwip: Return a partial or ENOBUFS from a non-blocking send. [merge of 9cdfb4e]
  mpy-cross:    +0 +0.000% 
   bare-arm:    +0 +0.000% 
minimal x86:    +0 +0.000% 
   unix x64:    +0 +0.000% standard
      stm32:    +0 +0.000% PYBV10
      esp32:    +0 +0.000% ESP32_GENERIC
     mimxrt:    +0 +0.000% TEENSY40
        rp2:   +16 +0.002% RPI_PICO_W
       samd:    +0 +0.000% ADAFRUIT_ITSYBITSY_M4_EXPRESS
  qemu rv32:    +0 +0.000% VIRT_RV32

Comment thread extmod/modlwip.c
if (socket->timeout == 0) {
if (write_len <= 1) {
MICROPY_PY_LWIP_EXIT
*_errno = MP_ENOBUFS;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Does it make sense to return EAGAIN like #19708?

@dpgeorge dpgeorge added the extmod Relates to extmod/ directory in source label Sep 19, 2026
Comment thread extmod/modlwip.c
// as a partial count. Only when not even one byte fits is it ENOBUFS -- a resource
// error distinct from EAGAIN, which POLLOUT (reading tcp_sndbuf, not the pool) would
// contradict into a busy-spin.
if (socket->timeout == 0) {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It might be a bit simpler in code logic to move this block down befor the call to poll_sockets() below.

@srgg

srgg commented Oct 2, 2026

Copy link
Copy Markdown
Author

@dpgeorge, thank you for the review.

There is an issue with ENOBUFS: mp_is_nonblocking_error accepts only EAGAIN/EWOULDBLOCK, so ENOBUFS would propagate out of asyncio's Stream.drain() and fail the mbedtls send callback. Therefore:

  • switched to EAGAIN; the check now sits where poll_sockets() was.
  • the halving retry is gone too: with Nagle on, tcp_pbuf_prealloc allocates a full-MSS pbuf regardless of length, so it only reduced the segment count.

Prior art, read from source: lwIP's own netconn returns ERR_WOULDBLOCK from lwip_netconn_do_writemore on ERR_MEM, and Linux returns EAGAIN from sk_stream_wait_memory; both then gate socket writability with a per-socket rule (NETCONN_FLAG_CHECK_WRITESPACE low-water marks, sk_stream_moderate_sndbuf), which cannot see memory held by another socket.

PR reworked, three commits:

  1. A non-blocking send blocked on ERR_MEM: the loop now returns EAGAIN for a non-blocking socket; nothing is queued, and the caller retries. Previously, write() blocked up to 10 s and raised ENOMEM.

  2. The socket timeout did not cover the whole send: one start tick now covers both the tcp_sndbuf == 0 wait and the ERR_MEM loop; ETIMEDOUT is returned at the timeout. The 10 s limit applies only to a socket with no timeout. Previously, the loop ended only at 10 s with ENOMEM, and each wait used its own clock, so a send could take about 2× the timeout.

  3. POLLOUT reported writable after that EAGAIN: it now requires one segment's memory: a pbuf of mss bytes, one tcp_seg tried and released, and a free slot in snd_queuelen. The send retries once at a single MSS before returning EAGAIN. Previously, POLLOUT only checked tcp_sndbuf > 0, so the next write() failed again. Cost: one mem_malloc/mem_free and one memp_malloc/memp_free per poll on a socket that has seen ERR_MEM; nothing on any other poll.

Tested on an OpenMV RT1062 (cyw43 Wi-Fi), v1.28-based tree with the same change on its mp_hal_delay_ms(50) loop:

  • EAGAIN return: one reader throttled to 20 KB/s, 320 s: worst write() 1553086 us and 54 events over 100 ms before; 1264–1316 us and 0 events after, five runs.
  • Socket timeout: two readers oversubscribing the 16 KB lwIP heap, settimeout(0.5), 240 s: 390 ETIMEDOUT, 0 ENOMEM.
  • POLLOUT gate: same two readers, a non-blocking socket polling POLLOUT after each EAGAIN, 240 s: the low-water rule answered writable and the next write failed on 1780 of 3509 polls; the allocation gate on 0 of 3016, with 3294 chunks written against 1566.

@srgg
srgg marked this pull request as draft October 2, 2026 16:05
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

extmod Relates to extmod/ directory in source

Projects

None yet

Development

Successfully merging this pull request may close these issues.

modlwip: non-blocking socket blocks in send when out of memory

2 participants