Conversation
721e159 to
03d5faa
Compare
rbradford
left a comment
There was a problem hiding this comment.
@likebreath / @sboeuf Need to review as VFIO experts.
| #[cfg(test)] | ||
| mod tests { | ||
| use super::*; | ||
|
|
||
| #[test] | ||
| fn uapi_layout() { | ||
| assert_eq!(VFIO_DEVICE_FEATURE(), 15221); | ||
| assert_eq!(IOMMU_IOAS_MAP_FILE(), 15247); | ||
| assert_eq!(size_of::<VfioDeviceFeatureDmaBuf>(), 40); | ||
| } | ||
| } |
There was a problem hiding this comment.
It only compared the hand-written ioctl numbers and struct size with the UAPI, and a mismatch would show up as a failing probe anyway. Removed.
03d5faa to
381875c
Compare
With vfio_p2p_dma, the default, the BARs of VFIO devices are mapped into the IOMMU so that devices can DMA to each other. On an iommufd this goes through IOMMU_IOAS_MAP, which pins pages with GUP and fails with EFAULT on the VM_PFNMAP mapping of a BAR, so a VM with a VFIO device does not boot with iommufd=on unless vfio_p2p_dma=off. Since Linux 6.19, vfio-pci can export a BAR as a dma-buf (VFIO_DEVICE_FEATURE_DMA_BUF) and iommufd can map one (IOMMU_IOAS_MAP_FILE). Map the BARs that way on an iommufd. When the kernel or the driver of a device cannot export BARs, log a warning and leave P2P DMA to that device off, as QEMU does. A device in an x_nv_gpudirect_clique fails to be added instead, as the clique tells the guest driver that P2P DMA works. vfio-pci revokes the dma-bufs of a device that stops decoding memory, and iommufd does not restore a revoked mapping. So unmap the BARs when the guest clears PCI_COMMAND_MEMORY, which Linux does while sizing BARs, and map them again when it sets it. Writes to PMCSR, and to PCIe Device Control or AF Control, which can start an FLR, may revoke them without a change to PCI_COMMAND: map the BARs afresh after those. vCPU threads write the config space, so their seccomp filter gets VFIO_DEVICE_FEATURE and IOMMU_IOAS_MAP_FILE, and the VMM thread gets IOMMU_IOAS_MAP_FILE. VFIO_DEVICE_FEATURE_DMA_BUF is not in vfio-bindings yet and is defined in pci/src/vfio_dmabuf.rs. Fixes: cloud-hypervisor#8895 Signed-off-by: Max Makarov <maxpain@linux.com> Assisted-by: Claude:Opus-5.5
381875c to
961f0d4
Compare
likebreath
left a comment
There was a problem hiding this comment.
Thanks for the contribution. We need to get proper support and abstraction done from iommufd and vfio crates first. Also, this is related to the ongoing GB support work. Let's align on a patch forward in #8895 first. Convert this to a draft for now.
| @@ -0,0 +1,127 @@ | |||
| // Copyright © 2026 Cloud Hypervisor Authors | |||
There was a problem hiding this comment.
This whole file and any direct usage of kernel's iommufd and vfio uAPIs should be pushed down to the iommufd and vfio crate.
I believe @yamahata already has these patches from his fork. Please feel free to reach out and collaborate.
| fn platform_cfg(iommufd: bool) -> String { | ||
| if iommufd { | ||
| "iommufd=on,vfio_p2p_dma=off".to_string() | ||
| "iommufd=on".to_string() |
There was a problem hiding this comment.
This would break the integration test. Our CI machine is still running old kernel that does not have dma-buf support.
Fixes #8895.
With
--platform iommufd=onand the defaultvfio_p2p_dma=on, a VM with a VFIO device does not boot: the BARs go throughIOMMU_IOAS_MAP, which pins pages with GUP, and GUP fails withEFAULTon theVM_PFNMAPmapping of a BAR. Since Linux 6.19 vfio-pci can export a BAR as a dma-buf (VFIO_DEVICE_FEATURE_DMA_BUF) and iommufd can map one (IOMMU_IOAS_MAP_FILE). This maps the BARs that way on an iommufd.VFIO_DEVICE_FEATURE_PROBE. Before 6.19 (ENOTTY), or with a driver that does not export BARs such as mlx5_vfio_pci (EOPNOTSUPP), the device gets a warning and no P2P DMA, like QEMU does for BARs it cannot map. A device withx_nv_gpudirect_cliquefails to be added instead, since the clique tells the guest driver that P2P DMA works.iopt_revoke_notify()). The mappings followPCI_COMMAND_MEMORY, and are redone after writes to PMCSR, PCIe Device Control and AF Control.VFIO_DEVICE_FEATUREandIOMMU_IOAS_MAP_FILEfor vCPU threads,IOMMU_IOAS_MAP_FILEfor the VMM thread.VFIO_DEVICE_FEATURE_DMA_BUFyet, so it is defined inpci/src/vfio_dmabuf.rstogether with the two ioctls. I can move them to rust-vmm/vfio and cloud-hypervisor/iommufd first if you prefer (question 1 in Support VFIO dma-buf for peer-to-peer DMA #8895).vfio_p2p_dma=off.Not covered: devices behind the virtual IOMMU, which still map BARs with
IOMMU_IOAS_MAP, and the device reset on the migration failure path, which is not followed by a remap (vfio-pci, the driver that exports dma-bufs here, has no migration support).Testing, with this change unmodified on top of our v54 build: Linux 7.2.6 host (2x EPYC 9575F), a VM with 8x RTX 6000D on vfio-pci and a ConnectX-7 VF on mlx5_vfio_pci, all passed by fd on iommufd.
EOPNOTSUPPwarning and works without P2P; the GPUs log nothing.IOMMU_IOAS_MAP_FILEfails withENODEVwhile in D3hot and succeeds on the D0 write), and peer copies to and from that GPU work once the driver is bound again.I could not run the NVIDIA integration tests.
LLM use: written with Claude Opus 5.5, see
Assisted-byin the commit.