Home / Alt manpages / perf-stat(1)

  • perf-stat(1)
  • User command
  • linux

Measure a Command with perf stat Without Misreading the Counters

You will finish with a repeatable way to measure a command's elapsed time and hardware events, select events deliberately, and save results for later comparison. The examples follow the perf-stat(1) manpage installed with linux-tools-common 6.8.0-142.142.

Allow 15 minutes for a first measurement, plus time to make the workload representative. You need the perf command, a shell and a command that is safe to run repeatedly. This guide observes a workload. It does not change kernel settings, but system-wide measurements can expose activity from other users and can be restricted by the host.

Local prerequisite check: run this before following the examples:

$ perf --version

On the machine used for this guide, /usr/bin/perf reports that matching tools for kernel 6.8.0-139-generic are unavailable. Install or select the matching distribution package before expecting a measurement. Do not treat a missing matching executable as a counter or workload failure.

1. Measure one command

Start with a short, deterministic command. Put -- before a command whose options might otherwise be read as perf stat options:

$ perf stat -- sh -c 'for n in 1 2 3 4 5; do :; done'

 Performance counter stats for 'sh -c for n in 1 2 3 4 5; do :; done':

       ...      cycles
       ...      instructions
       ... seconds time elapsed

       ...

The exact events and numbers depend on the processor and the installed perf build. The useful first result is a baseline, not a universal benchmark. The command runs as a child of perf, so include the complete workload in the measured command and keep setup outside it.

Checkpoint: repeat the command once. If elapsed time or counts vary widely, investigate the workload, CPU frequency, background activity and event availability before drawing a conclusion.

2. Choose events instead of accepting the default

List events supported by this host, then choose names from that output:

$ perf list
$ perf stat -e cycles,instructions -- ./your-program --input ./sample.dat

-e or --event= accepts a symbolic event name. Common names such as cycles and instructions are not guaranteed to have identical availability or meaning across processor families. Use the local perf list, not a command copied from a different machine.

Events may include modifiers, such as cpu-cycles:p, and the manpage also supports raw events and PMU-specific forms. Avoid raw encodings until you have checked this machine's PMU documentation and /sys/bus/event_source/devices/. A raw event copied from another CPU can fail or measure the wrong thing.

If perf reports that an event could not be opened, add -v to show counter-open errors. Do not silently replace a failed event with a different one and compare the results as though nothing changed.

3. Repeat the workload to see noise

Use -r or --repeat when a single run is not enough:

$ perf stat --repeat 5 -- ./your-program --input ./sample.dat

Perf repeats the command and prints an average and standard deviation. The documented maximum is 100 repetitions. --repeat 0 means repeat forever, so do not use it in a copy-and-paste benchmark unless you also have an explicit stop plan.

For a quick view of elapsed-time variation without starting hardware counters, use the null mode:

$ perf stat --null --repeat 5 --table -- ./your-program --input ./sample.dat

--null starts no counters. It is useful for wall-clock timing and for estimating perf stat overhead. The table shows individual measurements and a final result. For short commands, process startup and scheduling can dominate the number you hoped to measure, so use a workload long enough to represent the real operation.

4. Make output safe to parse

Perf writes its normal statistics to standard error. Redirect that stream when collecting a result, and use a field separator when another program will parse it:

$ perf stat -x , -e cycles,instructions -- ./your-program --input ./sample.dat 2> perf-stat.csv
$ sed -n '1,6p' perf-stat.csv

The separator form is CSV-style rather than a promise that every field is suitable for every CSV parser. Locale-aware number formatting is enabled by default, so large values may contain thousands separators. Use --no-big-num when a consumer needs ungrouped numbers, or set the corresponding perf config stat.big-num=false configuration deliberately.

For a machine-readable format, the manpage documents -j for JSON output. With interval output, JSON can include timestamps; with per-core, per-socket, per-die, per-node or per-thread aggregation, it can include the corresponding identifier. Pin the output format and perf version in an automated measurement job rather than assuming that every host exposes the same event names.

5. Measure over time or after startup

For a running interval, print counter deltas every second while a command runs:

$ perf stat -I 1000 -e cycles -- sleep 5

The interval is in milliseconds, with a documented minimum of 1 ms. Very short intervals add overhead and can make the measurement less representative. The manpage warns that the overhead percentage can be high below 100 ms.

If startup is not part of the question, delay collection:

$ perf stat --delay 2000 -e cycles,instructions -- ./your-program --input ./sample.dat

The delay is in milliseconds. This filters the first two seconds after the program starts, so state the delay in any report. It does not make a cold-cache measurement equivalent to a warm-cache measurement; it merely changes which part of execution is counted.

6. Understand aggregation and permissions

Without a process or thread target, the manpage describes collection across all CPUs as the default. Use -a explicitly when you mean system-wide collection, and combine it with -C 0,1 or a range such as -C 0-2 when only selected CPUs matter:

$ sudo perf stat -a -C 0-2 -e cycles -- sleep 5

This is the first example requiring elevated privileges in the guide because system-wide access is commonly restricted by the kernel and host policy. It also measures other work running on those CPUs. Ask for permission before using it on a shared machine, and do not collect system-wide counters when a per-process measurement answers the question.

By default, counts from multiple monitored CPUs can be aggregated. Add -A or --no-aggr when the per-CPU distribution is the result you need. For other layouts, the manpage provides --per-core, --per-socket, --per-die, --per-node and --per-cache; these are system-wide modes and need -a.

A count may be scaled when the hardware cannot run all requested events at once. The output can show a percentage indicating how long an event ran compared with how long it was enabled. A low percentage means the result has less direct coverage. Reduce the event list, repeat separate measurements, or report the limitation. Do not use --no-scale to hide multiplexing; it disables normal scaling and can make the raw number misleading.

7. Save and report a measurement

When a run must be reviewed later, record stat data and report it separately:

$ perf stat record -o perf-stat.data -- ./your-program --input ./sample.dat
$ perf stat report -i perf-stat.data

Keep the command line, host CPU model, kernel, perf version, event list, repetition count and relevant workload conditions next to the result. A counter value without those details is difficult to reproduce and easy to overinterpret.

The record command writes a file. If it is sensitive, restrict its permissions and remove it using your normal approved data-retention process after review. Do not upload perf data from a production host until you have checked what the workload and environment could reveal.

Common traps

  • A successful command does not prove that every requested event was measured continuously. Read the running-time percentage and any warnings.
  • CPU migrations, background activity and frequency changes can move results. Repeat runs and keep the workload conditions consistent.
  • -C selects CPUs, but it does not by itself turn on system-wide monitoring. The manpage says -a is still needed for that mode.
  • --metric-only prints computed metrics without raw values and is not supported with --per-thread. Do not use it when the raw event counts are needed for diagnosis.
  • Elevated privileges can solve an access restriction, but they do not make an unavailable PMU event valid and do not remove the privacy boundary of system-wide collection.

Done means

  • perf --version works with a matching executable for the running kernel.
  • You selected events from this host's perf list output.
  • You repeated the workload or used a null run when variation or timing overhead mattered.
  • You recorded event names, scaling or multiplexing warnings, host details and workload conditions.
  • Any system-wide run was deliberate, authorised and clearly labelled as measuring other activity too.