Profile Arm Memory Latency with perf and SPE
By the end of this guide you will have a perf.data recording from Arm's Statistical Profiling Extension (SPE), a report of the decoded instruction samples, and a way to inspect memory access details. The examples use the perf command-line interface supplied by the linux-tools-common package.
The route
Jump straight to the step you need, or tick off Done means at the end.
Allow 10 to 20 minutes for a first run. You need an Arm CPU with SPE support, a kernel with the ARM_SPE_PMU option enabled, and a workload that can be run repeatedly. The machine used for this guide is not suitable for a live capture: it is x86_64, and its installed perf reports that matching kernel tools are missing. Treat that as a useful reminder to check the target host first.
Checkpoint 1: confirm the target
Run these ordinary, read-only checks on the Arm host:
uname -m
perf --version
ls -ld /sys/bus/event_source/devices/arm_spe
You want an Arm machine, a working perf binary, and an arm_spe PMU directory. The local manpage is dated 1 September 2026. The installed wrapper here identifies itself as perf but warns that perf is not installed for kernel 6.8.0-139, so a version-looking command alone does not prove that recording will work.
If the PMU directory is absent, do not keep changing event syntax. The likely causes are a kernel without the SPE driver, a module that is not loaded, a virtual machine without exposed SPE hardware, or a kernel configuration that needs page table isolation disabled. The manpage specifically identifies kpti=off as a possible boot parameter and says the kernel will print profiling buffer inaccessible when that is required. Changing a boot command line affects the next boot and can reduce a security mitigation, so involve the machine owner and keep a way to restore the previous kernel parameters.
Checkpoint 2: record a small baseline
Start with the simplest capture. Replace ./mybench with a real executable and its arguments:
perf record -e arm_spe// -- ./mybench
This writes raw SPE data to perf.data in the current directory. Recording normally needs access to the performance monitoring interface; if the command is rejected by the system policy, ask an administrator about the permitted perf_event_paranoid setting or run the capture with the required privilege. Use sudo only when your host policy requires it, and remember that a root-created output file may not be writable by your normal account.
For a first capture, raise the sample period rather than collecting as aggressively as possible:
perf record -c 100000 -e arm_spe// -- ./mybench
The manpage says the period is programmed as the SPE interval and recommends a higher value because the hardware-derived minimum is used when no interval is specified. The exact useful value depends on the CPU and workload. A large output file is a practical sign that you need a higher interval. SPE is statistical, so a short run or a very sparse workload can legitimately produce little data.
Checkpoint 3: narrow the records
SPE configuration parameters go between the two slashes and are separated by commas. Capture loads with at least 10 units of latency like this:
perf record -c 100000 \
-e arm_spe/load_filter=1,min_latency=10/ -- ./mybench
min_latency is the total latency measured from the point at which sampling started on the instruction. It is not merely the time spent executing the load. Other useful filters are store_filter=1 for stores and branch_filter=1 for branches. These filters change what is retained; they do not turn SPE into a complete trace.
To retain only selected events, use the event mask documented by the manpage. Bit 1 is instruction retired, bit 3 is an L1 data-cache refill, bit 5 is a TLB refill, bit 7 is a mispredict, and bit 11 is a misaligned access. For retired instructions:
perf record -c 100000 \
-e arm_spe/event_filter=2/ -- ./mybench
For mispredicted branches, the corresponding example is:
perf record -c 100000 \
-e arm_spe/event_filter=0x80/ -- ./mybench
Leave jitter=1 enabled unless you have a reason to remove the pseudo-random interval perturbation. It helps avoid resonance between a regular workload and a regular sampling interval.
Checkpoint 4: decode and inspect
Recording does not decode SPE packets. Decoding happens when you open the file with perf report or perf script. First request one instruction sample per decoded instruction, without further downsampling:
perf report --itrace=i1i
The report may show groups such as arm_spe//, dummy:u, l1d-miss, tlb-access and memory. The first two are implementation details and are expected to be empty. The other groups are not necessarily unique samples: one instruction can have several associated events, so do not add the group counts as though they were independent observations.
To inspect load and store information attached to samples, use:
perf report --mem-mode
For a script-friendly view, use perf script on the same file. If you need to preserve the raw output while experimenting with reports, copy perf.data to a separate working directory first. The report commands are read-only; deleting the file is the only example in this guide that would remove captured evidence, so keep it until the result is recorded elsewhere.
Understand the limits before acting on a result
SPE samples include a program counter, timing, PMU events and, for loads and stores, data addresses, cache information and data origin. It provides precise attribution without tracing every operation, but it does not provide call-graph information. Results remain statistical and should be compared across repeatable runs.
Some implementations sample micro-operations rather than architectural instructions. An instruction expanded into two micro-operations is then twice as likely to enter the sample population. The manpage points to the sample_pop and inst_retired PMU events for estimating this coarse effect. Data-source meanings are also implementation defined because cache layouts differ between processors.
Physical address capture with pa_enable=1 and physical timestamping with pct_enable=1 require privilege. Avoid enabling them merely because they are available: physical addresses and timing details can expose more system information than a normal user-space profile needs.
Diagnose the common failures
- Cannot find PMU 'arm_spe'. Check the kernel driver, module loading, KPTI requirement, CPU architecture and whether a virtual machine hides SPE.
- Arm SPE CONTEXT packets not found. Root privilege is needed for context packets, which improve PID assignment for kernel samples. For a user-space-only profile, the manpage says this warning can be ignored.
- Too few samples. Run the workload longer or use a lower period, while watching output size and collision counts.
- Too many collisions. A new sample is dropped when the previous sampled operation has not finished. The
sample_collisionPMU event is a rough guide, but it counts collisions before filtering and is not an exact filtered-sample loss count. - Very large
perf.data. Increase the sampling interval and repeat the capture. Do not infer that a larger file means a more accurate profile.
Done means
arm_speis present on the Arm target and the kernel supports it.- A repeatable workload has produced
perf.datawith a deliberately chosen interval. perf report --itrace=i1iopens the file and its group counts are interpreted as potentially overlapping.perf report --mem-modehas been used when memory access details matter.- Collision, privilege, micro-operation weighting and implementation-defined data-source limits are recorded alongside any conclusion.