Skip to content

macOS TSO frames dropped when forwarded to MTU < 1500 network #2718

Description

@vormmt

Describe the bug

TCP uploads from macOS collapse when a container L3-forwards them into a network with MTU < 1500 (TSO frames dropped as "fragmentation needed")

Environment

  • macOS 15.7.9 (24G830), Apple M4 (Mac16,13), arm64
  • OrbStack 2.2.3 (2020300), commit c83556b0ef8f
  • Docker Engine 29.4.0, kernel 7.0.14-orbstack-00380-ga7e0a2dc9535
  • macOS net.inet.tcp.tso=1 (default). The macOS-side bridge of the Docker network (bridge10x, com.apple.NetworkSharing) reports options=63<RXCSUM,TXCSUM,TSO4,TSO6>.

Summary
A TCP connection from macOS to a container IP is DNAT-ed and routed by a Linux network namespace inside OrbStack into a network whose MTU is below 1500. That namespace drops the multi-segment (TSO) frames sent by macOS and answers with ICMP "fragmentation needed". Each individual segment fits the MTU, because the MSS is negotiated correctly, but the aggregated frames fail the IPv4 forwarding MTU check.

macOS lowers the path MTU as requested. This changes nothing, because its MSS already fits, so it keeps sending TSO frames. Every one of them is dropped and recovered only by the retransmission timer (~230 ms), and throughput collapses to ~6 KB/s. With MTU 1500 behind the hop, or with TSO disabled on macOS, the same path runs at full speed.

To Reproduce

Steps to reproduce
Run on macOS with OrbStack as the Docker context: ./repro.sh 1450, then ./repro.sh 1500 as a control.

#!/bin/bash
# macOS TCP -> container that L3-forwards (DNAT) into a network with the given MTU.
set -eu
MTU="${1:-1450}"
cleanup() {
  docker rm -f repro-router repro-sink >/dev/null 2>&1 || true
  docker network rm repro-inner repro-outer >/dev/null 2>&1 || true
}
trap cleanup EXIT
cleanup
docker network create repro-outer >/dev/null
docker network create --opt com.docker.network.driver.mtu="$MTU" repro-inner >/dev/null
docker run -d --rm --name repro-sink --network repro-inner busybox:1.36 \
  sh -c 'while true; do nc -l -p 9000 >/dev/null; done' >/dev/null
# Router: any image with iptables; rancher/k3s ships it in /bin/aux.
docker run -d --rm --name repro-router --network repro-outer --cap-add NET_ADMIN \
  --sysctl net.ipv4.ip_forward=1 --entrypoint /bin/sh rancher/k3s:v1.33.6-k3s1 -c 'sleep 3600' >/dev/null
docker network connect repro-inner repro-router
SINK_IP=$(docker inspect -f '{{(index .NetworkSettings.Networks "repro-inner").IPAddress}}' repro-sink)
ROUTER_IP=$(docker inspect -f '{{(index .NetworkSettings.Networks "repro-outer").IPAddress}}' repro-router)
docker exec repro-router /bin/aux/iptables -t nat -A PREROUTING -p tcp --dport 9000 -j DNAT --to-destination "$SINK_IP:9000"
docker exec repro-router /bin/aux/iptables -t nat -A POSTROUTING -o eth1 -j MASQUERADE
fragfails() { docker exec repro-router awk '/^Ip: [0-9]/{print $19}' /proc/net/snmp; }
sleep 10                                   # let the host learn the new container IP
ping -c 2 "$ROUTER_IP" >/dev/null
echo "inner_mtu=$MTU macos_tso=$(sysctl -n net.inet.tcp.tso) router=$ROUTER_IP sink=$SINK_IP"
echo "router FragFails before: $(fragfails)"
python3 -u - "$ROUTER_IP" <<'EOF'
import socket, sys, time
s = socket.create_connection((sys.argv[1], 9000), timeout=5)
print("mss", s.getsockopt(socket.IPPROTO_TCP, socket.TCP_MAXSEG))
s.settimeout(1)
buf, sent, t0, deadline = b"\0" * 65536, 0, time.time(), time.time() + 30
done = False
while time.time() < deadline and not done:
    try:
        if sent < 2 * 1024 * 1024:
            sent += s.send(buf)
        else:
            s.shutdown(socket.SHUT_WR)
            s.settimeout(max(0.1, deadline - time.time()))
            s.recv(1)
            done = True
    except socket.timeout:
        pass
print(f"{'delivered' if done else 'NOT delivered'}: {sent/1024:.0f} KiB accepted by macOS in {time.time()-t0:.1f}s")
s.close()
EOF
echo "router FragFails after:  $(fragfails)"

Actual results

MTU behind the forwarding hop MSS 2 MiB upload within 30 s Ip: FragFails in router netns
1500 1448 delivered in < 0.1 s 0
1480 1428 stalls at ~6 KB/s 67 in 15 s
1450 1398 NOT delivered, 303 KiB accepted by macOS 129
1450 with net.inet.tcp.tso=0 on macOS 1398 full speed, ~1.7 MB/s sustained, 0.8% retransmits (measured on the real-world setup below) 0

Expected
Uploads should work at any MTU behind the forwarding hop, as they do when the hop has MTU 1500. The individual segments already fit.

Evidence

  • tcpdump -i bridge10x on macOS shows outgoing frames with TCP payload 2796 (2 × 1398) and 5592 (4 × 1398). It also shows incoming ICMP type 3 code 4 need to frag (mtu 1450) quoting exactly those frames (IP total length 2848 / 5644, DF set). In one 60 s capture there were 147 TSO frames and 147 ICMP errors, matched 1:1 by absolute TCP sequence number. No single-MSS segment was dropped.
  • Ip: FragFails and IcmpMsg: OutType3 in the forwarding namespace grow by the same amount. The OrbStack Docker namespace itself shows no drops.
  • After the ICMP errors, route -n get <dst> on macOS shows mtu 1450, yet the dropped frames keep coming.
  • Without an L3 hop (sink attached directly to an MTU-1450 network and reached from macOS over L2), there is no loss.
  • Hypothesis: the GSO segment size attached to TSO frames entering the VM is derived from the 1500-byte host MTU rather than from the flow's actual segment size. ip_forward() / skb_gso_validate_network_len() then rejects them whenever the egress MTU is below 1500. This is consistent with 1480 failing and 1500 passing; the GSO size itself cannot be observed from the macOS side.

Real-world impact
Seen with a single-node k3s cluster (flannel VXLAN, pod MTU 1450) in OrbStack, where a LoadBalancer service (MetalLB L2) forwards to a pod. Any macOS client that uploads more than a few segments at a time is affected: browsers, CLIs, long-lived gRPC streams. On a long-lived upload stream, 25–46% of bytes were retransmitted, goodput dropped to ~17 KB/s, and latency-sensitive messages multiplexed on that stream were delayed by 5–15 s or timed out.

Connections originated by a QEMU VM with user-mode networking (for example UTM "Emulated VLAN") are affected as well, because they are re-originated by macOS sockets.

Workarounds

  • sudo sysctl -w net.inet.tcp.tso=0 on macOS. This is global and does not persist across reboots.
  • Keep MTU 1500 behind any forwarding hop, for example flannel host-gw instead of vxlan on a single-node k3s.

Expected behavior

No response

Diagnostic report (REQUIRED)

OrbStack info:
Version: 2.2.3
Commit: c83556b0ef8f1ba9a33abbb194622b6b7a1c0307 (v2.2.3)

System info:
macOS: 15.7.9 (24G830)
CPU: arm64, 10 cores
CPU model: Apple M4
Model: Mac16,13
Memory: 16 GiB

Full report: https://orbstack.dev/_admin/diag/orbstack-diagreport_2026-09-29T17-18-45.179285Z.zip

Screenshots and additional context (optional)

No response

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    t/bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions