Home / Alt manpages / perf-mem(1)

  • perf-mem(1)
  • User command
  • linux

Profile Memory Loads and Stores with perf mem

You will record a command's sampled memory accesses into perf.data, then inspect the result with perf mem report. The workflow is read-only apart from creating the profiling data, and it gives you a useful view of where sampled loads and stores occur.

Allow about fifteen minutes for a first run. You need the perf program, the matching kernel performance tools, and a command that performs enough memory work to produce samples. The installed package here is linux-tools-common 6.8.0-142.142, but the executable is not available for the running 6.8.0-139 kernel. Check your own installation before starting:

$ command -v perf
$ perf --version

If the command is missing, install the kernel-specific tools through your normal distribution process. Do not treat a package containing the manpage as proof that the matching executable is installed.

1. Check the available memory events

Ask perf mem record which event selectors the local build and processor expose:

$ perf mem record -e list

The output is hardware-dependent. Keep a selector from that list if you need a specific event. Otherwise, omit -e and let perf mem use its normal event selection. An event that exists in the manual or on another machine is not automatically available on this processor.

Checkpoint: if this command reports that an event cannot be opened, stop there and use one of the selectors printed by the command. Changing event names at random will not fix a missing hardware facility.

2. Record a short, representative command

Run the workload after record. This example uses a harmless command, but replace it with the real operation you want to investigate:

$ perf mem record -- /usr/bin/sleep 5
[ perf record: Woken up 1 times to write data ]
[ perf record: Captured and wrote perf.data ]

The exact status lines vary by version and workload. The useful result is a new perf.data file in the current directory. The -- separates perf options from the command and its arguments, which prevents an option intended for the workload from being consumed by perf.

Recording may require elevated privileges when the kernel's performance-event policy restricts ordinary users. Try it unprivileged first. If the command reports a permissions or counter-access failure, rerun only the same read-only profile with sudo after checking your local policy:

$ sudo perf mem record -- /path/to/workload --safe-test-option

Do not profile secrets, authentication commands or other sensitive workloads unless you have a clear reason. Sample records can include instruction and data addresses, process names and symbol information.

3. Report the recorded samples

With perf.data still in the current directory, display the memory profile:

$ perf mem report

The report invokes the normal perf reporting view with memory-access fields. By default it includes both loads and stores. Use --type load or --type store when you need one operation class:

$ perf mem report --type load
$ perf mem report --type store

The short form is -t. The command's default is load,store, so a report that contains both kinds is expected rather than evidence that the filter was ignored.

If the data file is elsewhere, point the report at it explicitly:

$ perf mem report --input=/path/to/perf.data

Use the exact file produced by the recording run. Do not add --force just to silence an ownership check. That option disables ownership validation, so use it only when you understand why the file belongs to a different user and accepting that boundary is safe.

4. Inspect raw samples when a script needs them

For a parseable, one-sample-per-line view, ask the report to dump decoded raw samples:

$ perf mem report --dump-raw-samples
$ perf mem report --dump-raw-samples --field-separator='|'

The default separator is a space. Choose a separator that cannot occur in the fields you expect, and treat the output as version- and hardware-dependent data rather than a stable interchange format. Capture it for inspection before building automation around particular columns.

5. Limit CPUs only when you have a reason

By default, perf monitors all CPUs. Restrict collection with --cpu, using a comma-separated list or a range without spaces:

$ perf mem record --cpu=0,1 -- /path/to/workload
$ perf mem record --cpu=0-2 -- /path/to/workload

CPU numbering and availability are host-specific. Check the machine before choosing a list, and expect an error if a requested CPU is offline or outside the host's range:

$ nproc
$ lscpu

There is no undo step for CPU selection. It applies only to that recording process; rerunning without --cpu returns to the default of all CPUs.

6. Interpret the result without overclaiming

On Intel systems, the reported memory latency is use-latency. It includes pipeline queueing delays as well as memory-subsystem latency, so do not present it as a pure DRAM or cache access time.

On Arm64, perf mem uses Statistical Profiling Extension sampling. Hardware and kernel support are required, and the statistical method means that many real memory operations will not appear in the report. A short or empty report can therefore indicate sampling limits rather than an absence of memory activity.

Keep the generated perf.data until you have checked the report. If it contains sensitive addresses or names, protect it with the same care as other diagnostic output. Once it is no longer needed, remove that specific file using your normal retention process; do not delete a directory or a collection of unrelated profiling data by accident.

Done means

  • perf is installed and its event list was checked on the target machine.
  • A representative command completed under perf mem record and produced perf.data.
  • perf mem report displayed the samples, with --type used only when a load or store filter was needed.
  • Any privilege escalation was limited to the profiling command and justified by a permission failure.
  • Latency and Arm64 sampling results are being treated as statistical measurements, not complete traces or pure memory timings.