On Thu, 2026-10-01 at 09:37 +0300, Leon Romanovsky wrote:
On Wed, Sep 30, 2026 at 04:57:04PM +0200, Thomas Hellström wrote:
On Wed, 2026-09-30 at 17:32 +0300, Leon Romanovsky wrote:
On Wed, Sep 30, 2026 at 01:52:39PM +0200, Christian König wrote:
On 9/30/26 13:43, Leon Romanovsky wrote:
On Wed, Sep 30, 2026 at 10:34:56AM +0200, Christian König wrote:
On 9/30/26 10:18, Leon Romanovsky wrote:
...
> > At a minimum, exporters need to pass `p2pdma_provider`.
No, exactly that is a no-go. The neither the framework nor the importer should see the p2pdma_provider.
Only fully translated addresses where the DMA access should happen.
> > If I keep the “dma-buf: Let exporters hand out the P2PDMA > provider behind a > buffer” patch, I can move the P2P TLP types back into > `p2pdma.c` and export > only the function that indicates whether ATS is required. > > Is it ok?
What you can do is to forward declare enum pci_p2pdma_map_type and than pass that 1 to 1 from the exporter to the importer.
Unfortunately, neither suggestion applies to RDMA NICs. They need to know, before mapping addresses, whether to create the memory region with ATS enabled.
The design principle here is that the final location and access path of the data isn't determined when the buffer is created.
The importer first need to attach before it can query such information from the exporter.
In attach yes, this is why importer digs in dma_buf ops to get p2pdma_provide, however it is before addresses are known.
The importer needs a way to obtain device information from the exporter so that it can configure itself correctly.
That won't work with DMA-buf then, the exporter is completely opaque to the importer and that is for really good reasons.
Why in the world does the importer needs to know the information from the exporter before the mapping is created?
It is the exporter who decides how data is accessed by the importer and not the other way around.
There are several reasons:
- This is how DMA-BUF MRs are built in RDMA. In mlx5, they rely
on the ODP mechanism, which requires an MKEY to be created first. See commit 90da7dc8206a (“RDMA/mlx5: Support dma-buf based userspace memory region”). 2. P2P routing is a property of devices, not memory. It is known and remains stable. 3. See the VFIO TPH ST discussion, where the requirement to obtain the exporter’s P2P information in the importer was raised again.
Thanks
Returning again to Jason's series. Let's say we'd add just the mapping type infrastructure, converted users of pcie_p2pdma only to use that and then we'd have access to per-mapping-type data. This could actually be done as a prereq for this series and merged separately. It's a couple of patches only.
I afraid that you over optimistic about the amount of work.
I was thinking something like this https://gitlab.freedesktop.org/thomash/xe-vibe/-/commit/0ab1cefa5c1439701e2a...
Although I have a couple of review comments on the infrastructure, and it doesn't include re-negotiation at map-time.
Jason's match() and finish() callbacks could compute the interesting routes at attach time, perhaps even condesed to whether IOVA is used and whether ATS translated packages have a direct route (which is what mlx5 care about AFAICT). This information is kept outside core dma- buf and would be specific to the pcie_p2p mapping type (interconnect) only rather than having functions and callbacks bloating the core dma- buf structures.
Then exactly where the cross-subsystem match() and finish() implementations should live I figure remain up for discussion and guidance by Christoph?
Maybe I'm wrong, and everything will work out. However, given Christian's feedback to determine the mapping type when the mapping is established, this approach won't work for an RDMA exporter.
I'm not 100% clear as to whether Christian meant the mapping-type would be re-negotiated if an agreed mapping type failed, or whether we would just allow a transparent fallback to system memory dma-buffers? Christian?
In any case, IMHO the infrastructure must allow for an importer to rather fail a mapping if the original attach-agreed mapping type was severed at mapping time, and then you'd get the same behaviour as you are sketching now? Or you could chose to take an unlikely slowpath to perform whatever's necessary to accomodate the new mapping type, even if that includes having to re-fault already ODP-faulted memory. Pinning would also be an option, I figure.
Thanks, Thomas
As I mentioned, mlx5 uses on-demand paging (ODP), creating mappings in response to page faults. To handle these faults with reasonable performance, mlx5 must create a memory region (MR), with or without ATS, and it is needed to be created before first page fault.
I'm afraid we're going in circles, so I'll drop the dma-buf patches for now and focus on fixing only the PCI ATS flow.
Thanks
Thanks, Thomas