On 10/1/26 10:27, Leon Romanovsky wrote:
On Thu, Oct 01, 2026 at 09:49:31AM +0200, Christian König wrote:
On 10/1/26 08:37, Leon Romanovsky wrote:
On Wed, Sep 30, 2026 at 04:57:04PM +0200, Thomas Hellström wrote:
On Wed, 2026-09-30 at 17:32 +0300, Leon Romanovsky wrote:
...
Returning again to Jason's series. Let's say we'd add just the mapping type infrastructure, converted users of pcie_p2pdma only to use that and then we'd have access to per-mapping-type data. This could actually be done as a prereq for this series and merged separately. It's a couple of patches only.
I afraid that you over optimistic about the amount of work.
Yeah, agree. The proposal looked good but there will probably be quite a bunch of work.
Jason's match() and finish() callbacks could compute the interesting routes at attach time, perhaps even condesed to whether IOVA is used and whether ATS translated packages have a direct route (which is what mlx5 care about AFAICT). This information is kept outside core dma-buf and would be specific to the pcie_p2p mapping type (interconnect) only rather than having functions and callbacks bloating the core dma-buf structures.
Then exactly where the cross-subsystem match() and finish() implementations should live I figure remain up for discussion and guidance by Christoph?
Maybe I'm wrong, and everything will work out. However, given Christian's feedback to determine the mapping type when the mapping is established, this approach won't work for an RDMA exporter.
As I mentioned, mlx5 uses on-demand paging (ODP), creating mappings in response to page faults. To handle these faults with reasonable performance, mlx5 must create a memory region (MR), with or without ATS, and it is needed to be created before first page fault.
Yeah, I feared that you have something like that. At least for the current DMA-buf semantics that is not something which fits into that model.
Background is that system memory is usually seen as fallback which should always work and you can have really strange combination of use cases.
So what can happen is that you create a mapping and P2P is possible without ATS but then some other device attaches and we suddenly have to use ATS because the buffer is now in system memory.
The importer reports ATS support on a per-mapping basis. If the exporter receives this hint from deviceA but not deviceB, it prepares different addresses for the two mappings.
The `ats_per_mapping` flag is stored in the importer structure: https://lore.kernel.org/linux-rdma/20260928-fix-p2p-acs-v4-0-v8-20-404453b9c...
In our example, if the mlx5 and XE importers both attach to the same exporter, they receive different mappings.
That isn't sufficient.
It is perfectly possible that P2P is disabled later on because of another device attaching or simply resource constrains.
So that your initial mapping has ATS enabled and then you get a mapping with ATS disabled is perfectly possible and even trivially trigger able through uAPI with some exporters.
Supporting that is a must have, even if it's slow. The only alternative I can see is to pin things but as I said that also has some other down sides.
Regards, Christian.
Thanks