Home / Alt manpages / perf-arm-spe(1)

  • perf-arm-spe(1)
  • User command
  • linux

Profile Arm Memory Latency with perf and SPE

By the end of this guide you will have a perf.data recording from Arm's Statistical Profiling Extension (SPE), a report of the decoded instruction samples, and a way to inspect memory access details. The examples use the perf command-line interface supplied by the linux-tools-common package.

Allow 10 to 20 minutes for a first run. You need an Arm CPU with SPE support, a kernel with the ARM_SPE_PMU option enabled, and a workload that can be run repeatedly. The machine used for this guide is not suitable for a live capture: it is x86_64, and its installed perf reports that matching kernel tools are missing. Treat that as a useful reminder to check the target host first.

Checkpoint 1: confirm the target

Run these ordinary, read-only checks on the Arm host:

uname -m
perf --version
ls -ld /sys/bus/event_source/devices/arm_spe

You want an Arm machine, a working perf binary, and an arm_spe PMU directory. The local manpage is dated 1 September 2026. The installed wrapper here identifies itself as perf but warns that perf is not installed for kernel 6.8.0-139, so a version-looking command alone does not prove that recording will work.

If the PMU directory is absent, do not keep changing event syntax. The likely causes are a kernel without the SPE driver, a module that is not loaded, a virtual machine without exposed SPE hardware, or a kernel configuration that needs page table isolation disabled. The manpage specifically identifies kpti=off as a possible boot parameter and says the kernel will print profiling buffer inaccessible when that is required. Changing a boot command line affects the next boot and can reduce a security mitigation, so involve the machine owner and keep a way to restore the previous kernel parameters.

Checkpoint 2: record a small baseline

Start with the simplest capture. Replace ./mybench with a real executable and its arguments:

perf record -e arm_spe// -- ./mybench

This writes raw SPE data to perf.data in the current directory. Recording normally needs access to the performance monitoring interface; if the command is rejected by the system policy, ask an administrator about the permitted perf_event_paranoid setting or run the capture with the required privilege. Use sudo only when your host policy requires it, and remember that a root-created output file may not be writable by your normal account.

For a first capture, raise the sample period rather than collecting as aggressively as possible:

perf record -c 100000 -e arm_spe// -- ./mybench

The manpage says the period is programmed as the SPE interval and recommends a higher value because the hardware-derived minimum is used when no interval is specified. The exact useful value depends on the CPU and workload. A large output file is a practical sign that you need a higher interval. SPE is statistical, so a short run or a very sparse workload can legitimately produce little data.

Checkpoint 3: narrow the records

SPE configuration parameters go between the two slashes and are separated by commas. Capture loads with at least 10 units of latency like this:

perf record -c 100000 \
  -e arm_spe/load_filter=1,min_latency=10/ -- ./mybench

min_latency is the total latency measured from the point at which sampling started on the instruction. It is not merely the time spent executing the load. Other useful filters are store_filter=1 for stores and branch_filter=1 for branches. These filters change what is retained; they do not turn SPE into a complete trace.

To retain only selected events, use the event mask documented by the manpage. Bit 1 is instruction retired, bit 3 is an L1 data-cache refill, bit 5 is a TLB refill, bit 7 is a mispredict, and bit 11 is a misaligned access. For retired instructions:

perf record -c 100000 \
  -e arm_spe/event_filter=2/ -- ./mybench

For mispredicted branches, the corresponding example is:

perf record -c 100000 \
  -e arm_spe/event_filter=0x80/ -- ./mybench

Leave jitter=1 enabled unless you have a reason to remove the pseudo-random interval perturbation. It helps avoid resonance between a regular workload and a regular sampling interval.

Checkpoint 4: decode and inspect

Recording does not decode SPE packets. Decoding happens when you open the file with perf report or perf script. First request one instruction sample per decoded instruction, without further downsampling:

perf report --itrace=i1i

The report may show groups such as arm_spe//, dummy:u, l1d-miss, tlb-access and memory. The first two are implementation details and are expected to be empty. The other groups are not necessarily unique samples: one instruction can have several associated events, so do not add the group counts as though they were independent observations.

To inspect load and store information attached to samples, use:

perf report --mem-mode

For a script-friendly view, use perf script on the same file. If you need to preserve the raw output while experimenting with reports, copy perf.data to a separate working directory first. The report commands are read-only; deleting the file is the only example in this guide that would remove captured evidence, so keep it until the result is recorded elsewhere.

Understand the limits before acting on a result

SPE samples include a program counter, timing, PMU events and, for loads and stores, data addresses, cache information and data origin. It provides precise attribution without tracing every operation, but it does not provide call-graph information. Results remain statistical and should be compared across repeatable runs.

Some implementations sample micro-operations rather than architectural instructions. An instruction expanded into two micro-operations is then twice as likely to enter the sample population. The manpage points to the sample_pop and inst_retired PMU events for estimating this coarse effect. Data-source meanings are also implementation defined because cache layouts differ between processors.

Physical address capture with pa_enable=1 and physical timestamping with pct_enable=1 require privilege. Avoid enabling them merely because they are available: physical addresses and timing details can expose more system information than a normal user-space profile needs.

Diagnose the common failures

  • Cannot find PMU 'arm_spe'. Check the kernel driver, module loading, KPTI requirement, CPU architecture and whether a virtual machine hides SPE.
  • Arm SPE CONTEXT packets not found. Root privilege is needed for context packets, which improve PID assignment for kernel samples. For a user-space-only profile, the manpage says this warning can be ignored.
  • Too few samples. Run the workload longer or use a lower period, while watching output size and collision counts.
  • Too many collisions. A new sample is dropped when the previous sampled operation has not finished. The sample_collision PMU event is a rough guide, but it counts collisions before filtering and is not an exact filtered-sample loss count.
  • Very large perf.data. Increase the sampling interval and repeat the capture. Do not infer that a larger file means a more accurate profile.

Done means

  • arm_spe is present on the Arm target and the kernel supports it.
  • A repeatable workload has produced perf.data with a deliberately chosen interval.
  • perf report --itrace=i1i opens the file and its group counts are interpreted as potentially overlapping.
  • perf report --mem-mode has been used when memory access details matter.
  • Collision, privilege, micro-operation weighting and implementation-defined data-source limits are recorded alongside any conclusion.