Since the commit:
18049c8cff9 ("perf/aux: Allocate non-contiguous AUX pages by default")
it changed the AUX buffer allocator to allocate AUX pages page-by-page (order=0) unless a PMU explicitly asks for contiguous allocations via the capability flag PERF_PMU_CAP_AUX_PREFER_LARGE. The goal was to make AUX allocation more memory-friendly by default, because not all PMUs require physically contiguous AUX pages and large contiguous allocations can contribute to fragmentation on long-running systems.
However, Arm SPE and CoreSight/TRBE rely on page-table translation when writing trace data to memory. With page-by-page AUX allocation, a large AUX buffer is mapped with many small mappings. This increases TLB pressure, in practice this can increase trace-buffer latency due to table translation walks (TTW) and contribute to trace discontinuities.
This series restores large AUX allocation for Arm CoreSight and SPE by setting PERF_PMU_CAP_AUX_PREFER_LARGE.
This is intended to work together with the mm large-mapping series [1]. That series allows vmap() to map physically contiguous pages with larger granules. With this series, perf first tries larger-order AUX allocations, and the vmap() code can then create larger mappings for the contiguous chunks.
The fragmentation concern from commit 18049c8cff9c should not block this opt-in. PERF_PMU_CAP_AUX_PREFER_LARGE is a preference, not a hard requirement. The AUX allocator already falls back to smaller orders when a high-order allocation fails. So this series gives Arm trace PMUs the performance benefit when large chunks are available.
The comparison below uses a baseline that already includes the mm large-mapping series [1]. "Baseline" means that the mm series is applied but this Arm PMU series is not. "Large AUX" means the same kernel plus this series. Some configurations to mitigate noise during test:
1) The tests were run with CPU10 isolated with the kernel parameter "isolcpus=10". 2) CPU10 was used as the traced CPU, the PMU counter CPU, and the workload CPU. The perf control tasks were pinned to CPU2 so that they did not add extra work on CPU10. 3) Each test was run for 10 iterations, and the tables report the average counter values across those runs.
The results show that using larger AUX mappings reduces the TLB pressure. This is mainly visible in the refill events: CoreSight/TRBE shows a large drop in l2d_tlb_refill and a smaller reduction in l1d_tlb_refill, while SPE also reduces l2d_tlb_refill. The dtlb_walk event also drops in both tests, which shows fewer data TLB walks after the AUX buffer can be mapped with larger granules.
ETM sparse branch delay (cs_etm, AUX 1GB)
taskset -c 2 perf stat -C 10 -e cycles:u,instructions:u,dtlb_walk:u,l1d_tlb:u,l1d_tlb_refill:u,l2d_tlb_refill:u \ -- taskset -c 2 perf record -C 10 -m ,1G -e cs_etm// \ -- taskset -c 10 ./sparse_branch_delay.elf
| | Baseline | Large map | | | | Metric | Avg. | Avg. | Delta | Change | |----------------+-----------+-----------+------------+---------| | dtlb_walk | 72.8 | 63.9 | -8.9 | -12.23% | | l1d_tlb | 7,434.4 | 1,982.2 | -5,452.2 | -73.34% | | l1d_tlb_refill | 163.7 | 148.2 | -15.5 | -9.47% | | l2d_tlb_refill | 161,884.9 | 513.1 | -161,371.8 | -99.68% |
SPE dd memory copy (arm_spe, AUX 512MB)
taskset -c 2 perf stat -C 10 -e cycles:u,instructions:u,dtlb_walk:u,l1d_tlb:u,l1d_tlb_refill:u,l2d_tlb_refill:u \ -- taskset -c 2 perf record -C 10 -m ,512M -e arm_spe_0/ts_enable=1,pa_enable=1,period=64,min_latency=0/ \ -- taskset -c 10 dd if=/dev/zero of=/dev/shm/dd_mem_test bs=1M count=1024 status=progress
| | Baseline | Large map | | | | Metric | Avg. | Avg. | Delta | Change | |----------------+-----------+-----------+------------+---------| | dtlb_walk | 1,760.2 | 1,387.9 | -372.3 | -21.15% | | l1d_tlb | 257,312.4 | 251,460.9 | -5,851.5 | -2.27% | | l1d_tlb_refill | 15,921.9 | 15,933.6 | 11.7 | +0.07% | | l2d_tlb_refill | 4,285.0 | 2,796.5 | -1,488.5 | -34.74% |
Note that after setting PREFER_LARGE for CoreSight and SPE, the existing AUX trace drivers either prefer large pages or, in the case of Intel BTS/PT, use the stronger AUX_NO_SG constraint. We can refactor this later by either dropping PREFER_LARGE entirely or reversing the flag if a driver needs discrete pages. For now, keep PREFER_LARGE to preserve flexibility in the allocation policy.
[1] https://lore.kernel.org/linux-mm/20260715120813.3609949-1-jiangwen6@xiaomi.c...
Signed-off-by: Leo Yan leo.yan@arm.com --- Dev Jain (1): coresight: perf: Prefer large AUX mappings
Leo Yan (1): perf: arm_spe: Prefer large AUX mappings
drivers/hwtracing/coresight/coresight-etm-perf.c | 3 ++- drivers/perf/arm_spe_pmu.c | 3 ++- 2 files changed, 4 insertions(+), 2 deletions(-) --- base-commit: db2ddb87143519e20a95aa36c60b36107b736a58 change-id: 20260717-perf_aux_trace_large_granule-d9b30cc14b5a
Best regards,