Profile Memory Loads and Stores with perf mem
You will record a command's sampled memory accesses into perf.data, then inspect the result with perf mem report. The workflow is read-only apart from creating the profiling data, and it gives you a useful view of where sampled loads and stores occur.
The route
Jump straight to the step you need, or tick off Done means at the end.
Allow about fifteen minutes for a first run. You need the perf program, the matching kernel performance tools, and a command that performs enough memory work to produce samples. The installed package here is linux-tools-common 6.8.0-142.142, but the executable is not available for the running 6.8.0-139 kernel. Check your own installation before starting:
$ command -v perf
$ perf --version
If the command is missing, install the kernel-specific tools through your normal distribution process. Do not treat a package containing the manpage as proof that the matching executable is installed.
1. Check the available memory events
Ask perf mem record which event selectors the local build and processor expose:
$ perf mem record -e list
The output is hardware-dependent. Keep a selector from that list if you need a specific event. Otherwise, omit -e and let perf mem use its normal event selection. An event that exists in the manual or on another machine is not automatically available on this processor.
Checkpoint: if this command reports that an event cannot be opened, stop there and use one of the selectors printed by the command. Changing event names at random will not fix a missing hardware facility.
2. Record a short, representative command
Run the workload after record. This example uses a harmless command, but replace it with the real operation you want to investigate:
$ perf mem record -- /usr/bin/sleep 5
[ perf record: Woken up 1 times to write data ]
[ perf record: Captured and wrote perf.data ]
The exact status lines vary by version and workload. The useful result is a new perf.data file in the current directory. The -- separates perf options from the command and its arguments, which prevents an option intended for the workload from being consumed by perf.
Recording may require elevated privileges when the kernel's performance-event policy restricts ordinary users. Try it unprivileged first. If the command reports a permissions or counter-access failure, rerun only the same read-only profile with sudo after checking your local policy:
$ sudo perf mem record -- /path/to/workload --safe-test-option
Do not profile secrets, authentication commands or other sensitive workloads unless you have a clear reason. Sample records can include instruction and data addresses, process names and symbol information.
3. Report the recorded samples
With perf.data still in the current directory, display the memory profile:
$ perf mem report
The report invokes the normal perf reporting view with memory-access fields. By default it includes both loads and stores. Use --type load or --type store when you need one operation class:
$ perf mem report --type load
$ perf mem report --type store
The short form is -t. The command's default is load,store, so a report that contains both kinds is expected rather than evidence that the filter was ignored.
If the data file is elsewhere, point the report at it explicitly:
$ perf mem report --input=/path/to/perf.data
Use the exact file produced by the recording run. Do not add --force just to silence an ownership check. That option disables ownership validation, so use it only when you understand why the file belongs to a different user and accepting that boundary is safe.
4. Inspect raw samples when a script needs them
For a parseable, one-sample-per-line view, ask the report to dump decoded raw samples:
$ perf mem report --dump-raw-samples
$ perf mem report --dump-raw-samples --field-separator='|'
The default separator is a space. Choose a separator that cannot occur in the fields you expect, and treat the output as version- and hardware-dependent data rather than a stable interchange format. Capture it for inspection before building automation around particular columns.
5. Limit CPUs only when you have a reason
By default, perf monitors all CPUs. Restrict collection with --cpu, using a comma-separated list or a range without spaces:
$ perf mem record --cpu=0,1 -- /path/to/workload
$ perf mem record --cpu=0-2 -- /path/to/workload
CPU numbering and availability are host-specific. Check the machine before choosing a list, and expect an error if a requested CPU is offline or outside the host's range:
$ nproc
$ lscpu
There is no undo step for CPU selection. It applies only to that recording process; rerunning without --cpu returns to the default of all CPUs.
6. Interpret the result without overclaiming
On Intel systems, the reported memory latency is use-latency. It includes pipeline queueing delays as well as memory-subsystem latency, so do not present it as a pure DRAM or cache access time.
On Arm64, perf mem uses Statistical Profiling Extension sampling. Hardware and kernel support are required, and the statistical method means that many real memory operations will not appear in the report. A short or empty report can therefore indicate sampling limits rather than an absence of memory activity.
Keep the generated perf.data until you have checked the report. If it contains sensitive addresses or names, protect it with the same care as other diagnostic output. Once it is no longer needed, remove that specific file using your normal retention process; do not delete a directory or a collection of unrelated profiling data by accident.
Done means
perfis installed and its event list was checked on the target machine.- A representative command completed under
perf mem recordand producedperf.data. perf mem reportdisplayed the samples, with--typeused only when a load or store filter was needed.- Any privilege escalation was limited to the profiling command and justified by a permission failure.
- Latency and Arm64 sampling results are being treated as statistical measurements, not complete traces or pure memory timings.