On 10/1/26 08:37, Leon Romanovsky wrote:
On Wed, Sep 30, 2026 at 04:57:04PM +0200, Thomas Hellström wrote:
On Wed, 2026-09-30 at 17:32 +0300, Leon Romanovsky wrote:
...
Returning again to Jason's series. Let's say we'd add just the mapping type infrastructure, converted users of pcie_p2pdma only to use that and then we'd have access to per-mapping-type data. This could actually be done as a prereq for this series and merged separately. It's a couple of patches only.
I afraid that you over optimistic about the amount of work.
Yeah, agree. The proposal looked good but there will probably be quite a bunch of work.
Jason's match() and finish() callbacks could compute the interesting routes at attach time, perhaps even condesed to whether IOVA is used and whether ATS translated packages have a direct route (which is what mlx5 care about AFAICT). This information is kept outside core dma-buf and would be specific to the pcie_p2p mapping type (interconnect) only rather than having functions and callbacks bloating the core dma-buf structures.
Then exactly where the cross-subsystem match() and finish() implementations should live I figure remain up for discussion and guidance by Christoph?
Maybe I'm wrong, and everything will work out. However, given Christian's feedback to determine the mapping type when the mapping is established, this approach won't work for an RDMA exporter.
As I mentioned, mlx5 uses on-demand paging (ODP), creating mappings in response to page faults. To handle these faults with reasonable performance, mlx5 must create a memory region (MR), with or without ATS, and it is needed to be created before first page fault.
Yeah, I feared that you have something like that. At least for the current DMA-buf semantics that is not something which fits into that model.
Background is that system memory is usually seen as fallback which should always work and you can have really strange combination of use cases.
So what can happen is that you create a mapping and P2P is possible without ATS but then some other device attaches and we suddenly have to use ATS because the buffer is now in system memory.
When that is just a problem of quality of service, in other words it just takes long, then I would say it is irrelevant for production because that use case will most likely not happen in your environment.
It's just that if you don't support it somebody can trivially let your device run into a deny of service from userspace and that is something people try to avoid.
Only when you say that technically doesn't work at all then we need to find another solution. It is for example possible to call dma_buf_pin() to disable page faults, but that often also let exporters disable P2P so that is most likely problematic as well.
Regards, Christian.
I'm afraid we're going in circles, so I'll drop the dma-buf patches for now and focus on fixing only the PCI ATS flow.
Thanks
Thanks, Thomas