On 10/1/26 08:17, Leon Romanovsky wrote:
On Wed, Sep 30, 2026 at 04:52:20PM +0200, Christian König wrote:
On 9/30/26 16:32, Leon Romanovsky wrote:
On Wed, Sep 30, 2026 at 01:52:39PM +0200, Christian König wrote:
...
The importer needs a way to obtain device information from the exporter so that it can configure itself correctly.
That won't work with DMA-buf then, the exporter is completely opaque to the importer and that is for really good reasons.
Why in the world does the importer needs to know the information from the exporter before the mapping is created?
It is the exporter who decides how data is accessed by the importer and not the other way around.
There are several reasons:
- This is how DMA-BUF MRs are built in RDMA. In mlx5, they rely on the ODP mechanism, which requires an MKEY to be created first. See commit 90da7dc8206a (“RDMA/mlx5: Support dma-buf based userspace memory region”).
I know that patch, but I absolutely don't understand where the problem is.
- P2P routing is a property of devices, not memory. It is known and remains stable.
No, absolutely not. You are making completely incorrect assumptions how DMA-buf works.
Maybe, but I understand how the kernel, PCI, and P2P DMA work, and dma-buf seems to operate in a parallel reality that is not aligned with any of them.
Yes, but that behavior predates the P2P DMA work by over a decade.
You can't ignore how existing drivers work just because you want PCI P2P to work this way.
Again: What is exported and from which path is up to the exporter and only determined when you create a DMA-buf mapping!
Let's discuss the PCI P2P case, where `p2pdma_provider` is present. In this flow, the exporter has a PCI device, and the importer also has one when it attaches.
No, they don't.
The decision about how to construct the addresses is made at that point. Note that both the exporter and importer are PCI P2P devices, so the importer should make the request.
Again, no.
The decision where to place things is done when the first mapping is created and that is documented and full intentional behavior since the very first merged version in 2011.
In most cases you don't even have a single device which is the DMA-buf exporter.
How is it possible for exporter with p2pdma_provider?
Only while creating the mapping the underlying access path is finally determined. It can be that P2P is used, it can be that internal connections are used, it can be that the data is moved to system memory....
But we are talking about in-tree exporters: RDMA, VFIO e.t.c
That's why we have the distinction between attaching and importer and the importer creating a mapping.
- See the VFIO TPH ST discussion, where the requirement to obtain the exporter’s P2P information in the importer was raised again.
As far as I can see there isn't any. The TPH/ST information are just passed through from the exporter to the importer.
We could define a bit better what needs to come first the mapping or the TPH/ST query but for their use case that is actually irrelevant.
They need an access to p2pdma_provider too. https://lore.kernel.org/linux-rdma/CAH3zFs2rhg6-b-pVDps1V8LsM4Dn48t4J4LFVyKO...
No they don't. They just need the routing information the same way as you do it here.
And Zhiping reply actually makes it 100% clear were the misunderstanding is here:
Two properties are worth stating explicitly:
- The metadata and its callback are per-dmabuf, while routing is per
attachment, so this deliberately uses a conservative all-or-nothing gate across the current attachments. It may withhold TPH from a direct importer when another attachment is not BUS_ADDR, but it cannot return a tag while any current attachment has a non-direct route.
Yes that is fully correct.
An importer attaching later does not change an existing importer's route, and future queries reevaluate the current attachment list.
No, that is incorrect! An importer attaching later eventually *does* change the routing!
The final routing is only determined on the first mapping call.
In other word the semantics is like this:
attachmentA = dma_buf_attach(dma_buf, importerA); ... attachmentB = dma_buf_attach(dma_buf, importerB); ... ... dma_resv_lock(dma_buf->resv, NULL); ... pci_info = dma_buf_get_tph_st_and_routing(attachmentA); mappingA = dma_buf_map_attachment(attachmentA, ...); ... dma_resv_add_fence(dma_buf->resv, async_signal_preventing_unmap); dma_resv_unlock(dma_buf->resv);
So that PCI info including TPH, ST and routing is only valid as long as you hold the lock of the DMA-buf.
When another dma_buf_attach() call comes after the mapping is already created the invalidate mappings callback is called and eventually the buffer is moved to a completely different location.
So on the next mapping call you can get different TPH, ST and routing information and depending on the internal topology of the exporter eventually a different PCI device you can do P2P with.
What can be is that not all importers support the invalidate mappings callback and/or dma_buf-pin() is called, but then exporters usually don't allow PCI P2P in the first place or at least only under restrictive limits (like for example cgroups).
And yes that semantics has been the one of DMA-buf for 15 years now and yes we can't change any of that on existing exporters since that is uAPI.
I hope that I finally made it clear how things work here. I'm kind of running out of ideas how to explain that.
Regards, Christian.
Thanks
Regards, Christian.
Thanks