Conversation
RTR hard-coded max_dest_rd_atomic and RTS hard-coded max_rd_atomic to 16. Some RNICs advertise fewer, so ibv_modify_qp() rejects the request and the endpoint fails to come up. Reuse the device-capability clamping already done in updateGlobalConfig(): cache max_qp_init_rd_atom/max_qp_rd_atom in GlobalConfig, lower them to the value from ibv_query_device(), and apply them when building the QP attributes. Add rdma_atomic_cap_test, a hardware-free regression that drives updateGlobalConfig() with a synthetic ibv_device_attr. Fixes kvcache-ai#4309
|
I was looking at this exact code path after reading #4309, and found this PR already up — so switching to review instead of a competing patch. Clamping to the device capability is the right fix, thanks — reusing One design question. Smaller second point: a device that doesn't support single-sided READ reports Not blocking this PR either way — just something worth deciding before it becomes hard to change. |
@sususama Thanks for the detailed review — really helpful. On (1): You're right that a process-wide cap means the weakest NIC On (2): Good catch. I'll use max(1, ...) so a device reporting I'll push an update shortly. Feedback welcome. |
Address review feedback on the device-capability clamp. Move the caps out of process-wide GlobalConfig into RdmaContext, next to the other per-NIC properties (numLagPorts(), activeMTU()) that RdmaEndPoint already consumes, and derive them from the same ibv_query_device() result. A slow NIC on a heterogeneous host no longer caps every QP in the process. Log a warning when a device advertises less than the historical default of 16. Floor the programmed depth at 1: a device reporting max_qp_rd_atom = 0 has no single-sided READ support, and programming 0 would let the QP reach RTS and fail on every later READ. Programming 1 keeps ibv_modify_qp() as the loud failure point. Rewrite rdma_atomic_cap_test to cover the pure clamp helper: capability below the default, capability above it, and the floor at 1.
he-yufeng
left a comment
There was a problem hiding this comment.
Verified in the repo container (mc-dev:local): built from this head and ran rdma_atomic_cap_test, 3/3 pass.
Read the full diff line by line. The shape is right where it matters: both atomic depths are clamped (the RTS initiator depth at the handshake site and the RTR responder depth in doSetupConnection), the values come from the device attr already queried at context construction so there is no extra verbs call, per-NIC values keep a slow card from capping a fast one on heterogeneous hosts, and the floor-at-1 choice makes a device advertising 0 fail in modify_qp rather than mysteriously on every later READ. The clamp is capped at the historical 16, so capable NICs see no behavior change. The hardware-free test covers exactly the boundary that matters.
One note for reviewers, not a blocker: the fail-loud-on-zero case still cannot be exercised without an affected NIC, so the on-device confirmation from #4309's reporter remains the missing runtime datapoint. Everything else about the change is compile-level and covered by the new test.
Looks right to me.
Thanks for the thorough review and the container validation. Agreed on the note — the fail-loud-on-zero path needs an affected NIC to Thanks again! |
Description
Fixes #4309. The RTR transition hard-coded
max_dest_rd_atomic = 16and the RTS transition hard-codedmax_rd_atomic = 16. Some RNICs advertise fewer, soibv_modify_qp()returns an error and the endpoint never reaches RTS.Reuse the device-capability clamping already performed by
updateGlobalConfig(): cachemax_qp_init_rd_atom/max_qp_rd_atominGlobalConfig, lower them to the values reported byibv_query_device(), and apply them when building the QP attributes. This is the same pattern already used formax_wr,max_sge,max_cqe, etc.Module
mooncake-transfer-engine)Type of Change
How Has This Been Tested?
Test commands:
Test results:
rdma_atomic_cap_testis hardware-free and drivesupdateGlobalConfig()with a syntheticibv_device_attr(device cap 4/8 clamps the config; a more capable device keeps the configured 8/12). On the test host all mlx5 HCAs advertisemax_qp_rd_atom = 16, so runtime behavior is unchanged there and the sub-16 path is covered by the unit test instead. The regression sweep stays green.Checklist
./scripts/code_format.sh(git-clang-format on changed lines is clean)AI Assistance Disclosure