Home / Alt manpages / perf-record(1)

  • perf-record(1)
  • User command
  • linux

Record a Useful Linux Profile with perf record

You will finish with a perf.data file containing a performance profile for one command, plus a repeatable way to add call stacks when the first recording shows only function names. The examples follow the perf record manual installed with linux-tools-common version 6.8.0-142.142.

Allow 10 to 20 minutes for a first profile. You need a shell, the perf tools matching the running kernel, and a workload you can run safely. Recording is normally read-only from the workload's point of view, but the data file can contain command names, paths and stack information. Keep it out of shared or sensitive directories.

1. Check the tool and choose a workload

Start by checking the kernel and the event list:

$ uname -r
6.8.0-139-generic
$ perf list | head -n 20

You should see the running kernel release followed by available hardware and software events. The package and the executable must agree with the kernel closely enough for the required PMU support. On this host, /usr/bin/perf reports that the tools for kernel 6.8.0-139 are missing, even though the manual comes from linux-tools-common 6.8.0-142.142. Install or select the matching, already-approved kernel tools before treating a failed command as a workload problem.

Checkpoint: do not continue until perf list runs without the missing-tools warning. You can still read the syntax in this guide, but a recording cannot be validated on a host with no usable binary.

2. Record one command

Use an explicit output path and an event name. cycles is a common starting event, but event availability is processor-dependent, so choose a name shown by perf list on your machine:

$ perf record -e cycles -o ./perf.data -- ./path/to/workload ARGUMENT
[workload output appears here]
[perf finishes when the workload exits]

The command runs the workload and gathers samples without displaying a profile. The -- separates perf's options from the command and its options. If the workload itself begins with an option-like argument, keep the separator. The output file is the recording, not a text report.

Verify that the command produced the file:

$ test -s ./perf.data && printf '%s\n' 'perf.data was created'
perf.data was created

If your program writes its own files, run it in a disposable working directory. Do not replace a valuable existing recording without checking the output path first. A new recording at the same path changes the evidence you meant to compare.

3. Read the recording

The manual describes perf report as the later inspection command. Run it against the file you just created:

$ perf report -i ./perf.data

Expect an interactive report with sampled locations, subject to the symbols available on the machine. Exit it with q. If your environment has no interactive terminal, use the report command's help for the non-interactive output option supplied by that installed build rather than guessing at an option.

The busiest entries are where the samples landed, not automatically the functions that should be rewritten. First check that the workload and event are the ones you intended. A short run may produce too little evidence, and a profile of a mostly idle program can be dominated by startup or library code.

4. Add call graphs when names are not enough

Use -g to record call graphs for user and kernel space:

$ perf record -e cycles -g -o ./perf-callgraph.data -- ./path/to/workload ARGUMENT
$ test -s ./perf-callgraph.data && printf '%s\n' 'call-graph recording was created'
call-graph recording was created

The default user-space call-graph method is fp, which relies on frame pointers. If the workload was built with frame pointers omitted, those stacks can be wrong. Where the installed perf build provides suitable unwinding support, try DWARF and set an explicit stack dump size:

$ perf record -e cycles --call-graph dwarf,4096 -o ./perf-dwarf.data -- ./path/to/workload ARGUMENT

DWARF recording adds user stack data to samples and can increase overhead and file size. The manual gives 8192 bytes as its default stack dump size; the example chooses 4096 deliberately for a smaller trial. Compare the resulting report with the frame-pointer version instead of assuming one method is universally better.

5. Profile an existing process or the whole system

For a process that is already running, pass its PID:

$ perf record -p 12345 -e cycles -o ./pid-12345.data

Stop the recording with Ctrl-C after the behaviour you need has occurred, then verify the output file. Replace 12345 with a real PID. Do not attach to a production service casually: profiling can add overhead, and the resulting data may expose its command lines and code paths.

For all CPUs, use -a:

$ sudo perf record -a -e cycles -o ./system.data -- sleep 10
$ test -s ./system.data && printf '%s\n' 'system recording was created'
system recording was created

This captures system-wide activity for the ten-second window, not just your shell command. Elevated privilege may be required by the kernel's perf-event policy. Use sudo only when the command reports a permission failure, and remember that system-wide data includes other users' activity. Remove or protect the file after analysis according to your data-handling rules.

6. Diagnose the common traps

  • Event rejected: run perf list and choose an event actually available on this CPU. Do not copy a raw rN event from another processor.
  • Permission denied: try the smallest scope first, then check the host's perf-event security policy. -a, kernel events and some filters need more privilege than profiling your own command.
  • Empty or tiny data: make the workload run long enough to sample, and confirm that it really exercised the path of interest.
  • Missing or misleading stacks: use -g, check frame-pointer build settings, and try --call-graph dwarf only when the required unwinder is available.
  • Unexpected coverage: without an explicit target, perf record uses system-wide collection from all CPUs according to the manual. For a single command, put the command after --; use -p or -a when you mean to select an existing process or every CPU.

Done means

  • perf list works with tools suitable for the running kernel.
  • A deliberately named, non-empty perf.data file exists.
  • perf report opens that recording.
  • Call graphs were added only when their overhead and unwinding method were understood.
  • PID and system-wide recordings were treated as potentially sensitive and removed or protected after analysis.