On 9/17/26 15:56, Thomas Hellström wrote:
Hi
On Mon, 2026-09-14 at 14:22 +0300, Leon Romanovsky wrote:
From: Leon Romanovsky leonro@nvidia.com
mlx5_umem_needs_ats() enables ATS for any dma-buf whose caller asked for Relaxed Ordering, on the assumption that a switch in the path has CR, RR and DT all set. It also enables it for a buffer already mapped with the peer's bus addresses, which are not translatable at all.
P2PDMA has read the ACS controls, so ask it through dma_buf_p2pdma_map_type(): enable ATS only where the path is not routed directly as it stands, but would be for a Translated Request whose Completions carry Relaxed Ordering. Exporters that name no provider keep the old assumption, since their ACS settings remain hidden.
Signed-off-by: Leon Romanovsky leonro@nvidia.com
drivers/infiniband/hw/mlx5/mlx5_ib.h | 36 ++------------------------
drivers/infiniband/hw/mlx5/mr.c | 40 ++++++++++++++++++++++++++++++++++++ 2 files changed, 42 insertions(+), 34 deletions(-)
diff --git a/drivers/infiniband/hw/mlx5/mlx5_ib.h b/drivers/infiniband/hw/mlx5/mlx5_ib.h index e9ddf2e97a76..ab32742b2180 100644 --- a/drivers/infiniband/hw/mlx5/mlx5_ib.h +++ b/drivers/infiniband/hw/mlx5/mlx5_ib.h @@ -1646,40 +1646,8 @@ static inline bool rt_supported(int ts_cap) ts_cap == MLX5_TIMESTAMP_FORMAT_CAP_FREE_RUNNING_AND_REAL_TIME; } -/*
- PCI Peer to Peer is a trainwreck. If no switch is present then
things
- sometimes work, depending on the pci_distance_p2p logic for
excluding broken
- root complexes. However if a switch is present in the path, then
things get
- really ugly depending on how the switch is setup. This table
assumes that the
- root complex is strict and is validating that all req/reps are
matches
- perfectly - so any scenario where it sees only half the
transaction is a
- failure.
- CR/RR/DT ATS RO P2P
- 00X X X OK
- 010 X X fails (request is routed to root but root never
sees comp)
- 011 0 X fails (request is routed to root but root never
sees comp)
- 011 1 X OK
- 10X X 1 OK
- 101 X 0 fails (completion is routed to root but root
didn't see req)
- 110 X 0 SLOW
- 111 0 0 SLOW
- 111 1 0 fails (completion is routed to root but root
didn't see req)
- 111 1 1 OK
- Unfortunately we cannot reliably know if a switch is present or
what the
- CR/RR/DT ACS settings are, as in a VM that is all hidden. Assume
that
- CR/RR/DT is 111 if the ATS cap is enabled and follow the last
three rows.
- For now assume if the umem is a dma_buf then it is P2P.
- */
-static inline bool mlx5_umem_needs_ats(struct mlx5_ib_dev *dev,
struct ib_umem *umem, intaccess_flags) -{
- if (!MLX5_CAP_GEN(dev->mdev, ats) || !umem->is_dmabuf)
return false;- return access_flags & IB_ACCESS_RELAXED_ORDERING;
-} +bool mlx5_umem_needs_ats(struct mlx5_ib_dev *dev, struct ib_umem *umem,
int access_flags);int set_roce_addr(struct mlx5_ib_dev *dev, u32 port_num, unsigned int index, const union ib_gid *gid, diff --git a/drivers/infiniband/hw/mlx5/mr.c b/drivers/infiniband/hw/mlx5/mr.c index 00e13028762a..286f372e5b0c 100644 --- a/drivers/infiniband/hw/mlx5/mr.c +++ b/drivers/infiniband/hw/mlx5/mr.c @@ -38,6 +38,7 @@ #include <linux/export.h> #include <linux/delay.h> #include <linux/dma-buf.h> +#include <linux/dma-buf-mapping.h> #include <linux/dma-resv.h> #include <rdma/frmr_pools.h> #include <rdma/ib_umem_odp.h> @@ -47,6 +48,45 @@ #include "data_direct.h" #include "dmah.h" +MODULE_IMPORT_NS("DMA_BUF");
+bool mlx5_umem_needs_ats(struct mlx5_ib_dev *dev, struct ib_umem *umem,
int access_flags)+{
- struct dma_buf_attachment *attach;
- if (!MLX5_CAP_GEN(dev->mdev, ats) || !umem->is_dmabuf)
return false;- /*
* The Completer decides whether its Completions carryRelaxed
* Ordering, and only a Request that asked for it can expectthem to.
*/- if (!(access_flags & IB_ACCESS_RELAXED_ORDERING))
return false;- attach = to_ib_umem_dmabuf(umem)->attach;
- switch (dma_buf_p2pdma_map_type(attach, 0)) {
- case PCI_P2PDMA_MAP_NONE:
/* Nothing is known about the route, so fall back tothe bet. */
return true;- case PCI_P2PDMA_MAP_BUS_ADDR:
/** The path is routed directly already and isprogrammed with
* the peer's bus addresses. Those are nottranslatable, so
* ATS would be wrong as well as pointless.*/return false;- default:
break;- }
- return dma_buf_p2pdma_map_type(attach,
PCI_P2PDMA_TLP_TRANSLATED |PCI_P2PDMA_TLP_RELAXED_CPL) ==
- PCI_P2PDMA_MAP_BUS_ADDR;
+}
It looks like this works well when the device can choose whether to enable ATS per transaction.
However, at least for the Intel GPUs, ATS enablement is based on the PCIe-side enable bit.
Yeah, I can confirm as well that this is something really exotic.
I never heard of any other device who can enable/disable ATS on a transaction basis.
Additional to that the setup that ACS seems to be disabled for some bridges bit not all also sounds extremely uncommon.
This means that if p2pdma tells the dma-mapping layer to give an Xe device a bus address rather than an IOVA, things break, while pci_p2pdma_distance says everything is OK.
It looks like the infrastructure and solution added in this series is targeted at fixing this on the device side by conditionally enabling ATS. However I think we need to look also at having the computed routing assume untranslated transactions using IOVA rather than bus address.
That is, a flag to tell the topology check that some transactions *will* take the host-bridge path due to IOVA being used, and that the computations including pci_p2pdma_distance() need to check whether that is possible (checking whitelist etc.) and return the corresponding THRU_HOST_BRIDGE mapping type. Translated transactions taking a short- cut using the bus-address would then be hidden from the driver.
Whether that is best done as a parameter to these functions or perhaps as a flag in the PCI device, I'm not sure.
That sounds totally reasonable to me as well.
Regards, Christian.
Thanks, Thomas
static int mkey_max_umr_order(struct mlx5_ib_dev *dev) { if (MLX5_CAP_GEN(dev->mdev, umr_extended_translation_offset))