Skip to content

Post small RDMA messages inline - #3569

Open
alxrxs wants to merge 1 commit into
apache:masterfrom
alxrxs:rdma-inline-small-messages
Open

alxrxs wants to merge 1 commit into
apache:masterfrom
alxrxs:rdma-inline-small-messages

Conversation

@alxrxs

@alxrxs alxrxs commented Sep 28, 2026 •

Copy link
Copy Markdown

What problem does this PR solve?

Issue Number: none

Problem Summary:

RdmaEndpoint::CutFromIOBufList() posts every message by reference, so for a small RPC request or response the NIC first has to DMA-read the data from host memory. The QP requests no inline data (MAX_INLINE_DATA was never used and was commented out in #2876).

What is changed and the side effects?

Changed:

  • AllocateQp() requests 236 bytes of inline data, retries without it if the device refuses, and keeps the granted size in RdmaResource.
  • CutFromIOBufList() sets IBV_SEND_INLINE when a message fits that size and all of its blocks come from the RDMA block pool. User-registered memory is never inlined, since it may be GPU memory.

236 bytes is the largest Send whose inlined WQE fits the 256-byte BlueFlame buffer of current mlx5 NICs. Larger messages are posted as before.

Side effects:

  • Performance effects: removes a host memory read before each small message is sent. In a separate ping-pong microbenchmark that sends the same small RDMA write with and without IBV_SEND_INLINE, inlining cut one-way latency by about 0.43 us on Xeon 8480C hosts with ConnectX-7 and 0.89 us on EPYC 7742 hosts with ConnectX-6. I haven't measured brpc itself. With the default rdma_max_sge the send WQE does not grow, but with a small rdma_max_sge it can.

  • Breaking backward compatibility: no.


Check List:

Tested on soft-RoCE (rxe) with the rdma_performance server and an echo client: attachments from 0 to 300 KB compared byte for byte, messages up to 236 bytes went inline and larger ones by reference, and the same again with the QP creation retry forced. libbrpc.a builds with WITH_RDMA=ON.

Generated with Claude Code

CutFromIOBufList() posts every message by reference, so the NIC has to
read a small RPC from host memory before sending it. Request 236 bytes
of inline data at QP creation, retrying without it if the device
refuses, and inline a message that fits the granted size and comes
entirely from the RDMA block pool. 236 bytes is the largest Send whose
WQE fits the 256-byte BlueFlame buffer of current mlx5 NICs.

Generated-by: Claude Code (Claude Opus 5.5)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@alxrxs
alxrxs force-pushed the rdma-inline-small-messages branch from 91efa9b to 7da4d18 Compare October 2, 2026 12:56
@alxrxs
alxrxs marked this pull request as ready for review October 2, 2026 13:57
@chenBright
chenBright requested a balanced review from Copilot October 3, 2026 03:54

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

The new inline-selection and QP-fallback paths need automated regression coverage.

Review effort: Balanced
Findings: 2 Low severity

Open (2)
What changed in this PR

Adds inline posting for small RDMA messages to reduce DMA-read latency.

Changes:

  • Requests up to 236 bytes of inline QP capacity with fallback.
  • Inlines eligible pool-backed messages while excluding user-registered memory.
File Description
src/​brpc/​rdma/​rdma_endpoint.h Stores granted inline capacity.
src/​brpc/​rdma/​rdma_endpoint.cpp Negotiates capacity and selects inline sends.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment on lines +907 to +908
if (in_pool && this_len <= _resource->max_inline_data) {
wr.send_flags |= IBV_SEND_INLINE;
Comment on lines +1157 to +1160
if (qp == nullptr) {
// The device may not support inline data, try again without it
attr.cap.max_inline_data = 0;
qp = IbvCreateQp(GetRdmaPd(), &attr);
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants