Skip to content

pci, vmm: Map VFIO BARs for P2P DMA as dma-bufs on iommufd - #8973

Draft
maxpain wants to merge 1 commit into
cloud-hypervisor:mainfrom
maxpain:vfio-iommufd-dmabuf-p2p
Draft

maxpain wants to merge 1 commit into
cloud-hypervisor:mainfrom
maxpain:vfio-iommufd-dmabuf-p2p

Conversation

@maxpain

@maxpain maxpain commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Fixes #8895.

With --platform iommufd=on and the default vfio_p2p_dma=on, a VM with a VFIO device does not boot: the BARs go through IOMMU_IOAS_MAP, which pins pages with GUP, and GUP fails with EFAULT on the VM_PFNMAP mapping of a BAR. Since Linux 6.19 vfio-pci can export a BAR as a dma-buf (VFIO_DEVICE_FEATURE_DMA_BUF) and iommufd can map one (IOMMU_IOAS_MAP_FILE). This maps the BARs that way on an iommufd.

  • Support is probed once per device with VFIO_DEVICE_FEATURE_PROBE. Before 6.19 (ENOTTY), or with a driver that does not export BARs such as mlx5_vfio_pci (EOPNOTSUPP), the device gets a warning and no P2P DMA, like QEMU does for BARs it cannot map. A device with x_nv_gpudirect_clique fails to be added instead, since the clique tells the guest driver that P2P DMA works.
  • vfio-pci revokes the dma-bufs of a device that stops decoding memory, is reset or goes to D3hot, and iommufd does not restore a revoked mapping (iopt_revoke_notify()). The mappings follow PCI_COMMAND_MEMORY, and are redone after writes to PMCSR, PCIe Device Control and AF Control.
  • seccomp: VFIO_DEVICE_FEATURE and IOMMU_IOAS_MAP_FILE for vCPU threads, IOMMU_IOAS_MAP_FILE for the VMM thread.
  • vfio-bindings does not have VFIO_DEVICE_FEATURE_DMA_BUF yet, so it is defined in pci/src/vfio_dmabuf.rs together with the two ioctls. I can move them to rust-vmm/vfio and cloud-hypervisor/iommufd first if you prefer (question 1 in Support VFIO dma-buf for peer-to-peer DMA #8895).
  • The iommufd integration tests no longer set vfio_p2p_dma=off.

Not covered: devices behind the virtual IOMMU, which still map BARs with IOMMU_IOAS_MAP, and the device reset on the migration failure path, which is not followed by a remap (vfio-pci, the driver that exports dma-bufs here, has no migration support).

Testing, with this change unmodified on top of our v54 build: Linux 7.2.6 host (2x EPYC 9575F), a VM with 8x RTX 6000D on vfio-pci and a ConnectX-7 VF on mlx5_vfio_pci, all passed by fd on iommufd.

  • The VF logs the EOPNOTSUPP warning and works without P2P; the GPUs log nothing.
  • CUDA peer copies of 256 MiB between all 56 ordered GPU pairs, checked by reading back: no mismatch, up to 52 GB/s.
  • With one GPU unbound from the guest driver, D3hot and back to D0 through PMCSR, and an FLR through Device Control: bpftrace on the VMM shows the BARs unmapped and mapped again within the write itself (IOMMU_IOAS_MAP_FILE fails with ENODEV while in D3hot and succeeds on the D0 write), and peer copies to and from that GPU work once the driver is bound again.
  • The same peer copies after a guest reboot.

I could not run the NVIDIA integration tests.

LLM use: written with Claude Opus 5.5, see Assisted-by in the commit.

@maxpain
maxpain requested a review from a team as a code owner September 30, 2026 14:55
@maxpain
maxpain force-pushed the vfio-iommufd-dmabuf-p2p branch from 721e159 to 03d5faa Compare September 30, 2026 16:37

@rbradford rbradford left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@likebreath / @sboeuf Need to review as VFIO experts.

Comment thread pci/src/vfio_dmabuf.rs Outdated
Comment on lines +116 to +126
#[cfg(test)]
mod tests {
use super::*;

#[test]
fn uapi_layout() {
assert_eq!(VFIO_DEVICE_FEATURE(), 15221);
assert_eq!(IOMMU_IOAS_MAP_FILE(), 15247);
assert_eq!(size_of::<VfioDeviceFeatureDmaBuf>(), 40);
}
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What's the point of this ?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It only compared the hand-written ioctl numbers and struct size with the UAPI, and a mismatch would show up as a failing probe anyway. Removed.

@maxpain
maxpain force-pushed the vfio-iommufd-dmabuf-p2p branch from 03d5faa to 381875c Compare September 30, 2026 19:05
With vfio_p2p_dma, the default, the BARs of VFIO devices are mapped
into the IOMMU so that devices can DMA to each other. On an iommufd this
goes through IOMMU_IOAS_MAP, which pins pages with GUP and fails with
EFAULT on the VM_PFNMAP mapping of a BAR, so a VM with a VFIO device
does not boot with iommufd=on unless vfio_p2p_dma=off.

Since Linux 6.19, vfio-pci can export a BAR as a dma-buf
(VFIO_DEVICE_FEATURE_DMA_BUF) and iommufd can map one
(IOMMU_IOAS_MAP_FILE). Map the BARs that way on an iommufd. When the
kernel or the driver of a device cannot export BARs, log a warning and
leave P2P DMA to that device off, as QEMU does. A device in an
x_nv_gpudirect_clique fails to be added instead, as the clique tells
the guest driver that P2P DMA works.

vfio-pci revokes the dma-bufs of a device that stops decoding memory,
and iommufd does not restore a revoked mapping. So unmap the BARs when
the guest clears PCI_COMMAND_MEMORY, which Linux does while sizing BARs,
and map them again when it sets it. Writes to PMCSR, and to PCIe Device
Control or AF Control, which can start an FLR, may revoke them without
a change to PCI_COMMAND: map the BARs afresh after those.

vCPU threads write the config space, so their seccomp filter gets
VFIO_DEVICE_FEATURE and IOMMU_IOAS_MAP_FILE, and the VMM thread gets
IOMMU_IOAS_MAP_FILE. VFIO_DEVICE_FEATURE_DMA_BUF is not in vfio-bindings
yet and is defined in pci/src/vfio_dmabuf.rs.

Fixes: cloud-hypervisor#8895

Signed-off-by: Max Makarov <maxpain@linux.com>
Assisted-by: Claude:Opus-5.5
@maxpain
maxpain force-pushed the vfio-iommufd-dmabuf-p2p branch from 381875c to 961f0d4 Compare September 30, 2026 19:39

@likebreath likebreath left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the contribution. We need to get proper support and abstraction done from iommufd and vfio crates first. Also, this is related to the ongoing GB support work. Let's align on a patch forward in #8895 first. Convert this to a draft for now.

Comment thread pci/src/vfio_dmabuf.rs
@@ -0,0 +1,127 @@
// Copyright © 2026 Cloud Hypervisor Authors

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This whole file and any direct usage of kernel's iommufd and vfio uAPIs should be pushed down to the iommufd and vfio crate.

I believe @yamahata already has these patches from his fork. Please feel free to reach out and collaborate.

fn platform_cfg(iommufd: bool) -> String {
if iommufd {
"iommufd=on,vfio_p2p_dma=off".to_string()
"iommufd=on".to_string()

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This would break the integration test. Our CI machine is still running old kernel that does not have dma-buf support.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Support VFIO dma-buf for peer-to-peer DMA

3 participants