Home / Alt manpages / perf-bench(1)

  • perf-bench(1)
  • User command
  • linux

Measure Linux Scheduler and Memory Paths with perf bench

You will run a small, repeatable benchmark with perf bench, select output that suits either a person or a script, and vary one workload parameter at a time. The examples use the perf-bench(1) interface shipped by linux-tools-common version 6.8.0-142.142. The installed wrapper on this machine also reports that matching tools for kernel 6.8.0-139 are missing, so the commands are valid but may need the kernel-specific perf package before they can execute here.

Allow about fifteen minutes for a first comparison, plus time for repeated runs if the machine is busy. You need a shell, a matching perf binary, and permission to run ordinary user-space benchmarks. Nothing in this guide changes a service, cgroup, kernel setting or persistent file.

1. Check the installed tool before benchmarking

Start with read-only checks. These do not require elevated privileges:

$ command -v perf
/usr/bin/perf
$ dpkg-query -W -f='${Package} ${Version}\n' linux-tools-common
linux-tools-common 6.8.0-142.142
$ perf --version
WARNING: perf not found for kernel 6.8.0-139

You may need to install the following packages for this specific kernel:
    linux-tools-6.8.0-139-generic
    linux-cloud-tools-6.8.0-139-generic

On a host with the matching package installed, the last command should identify a working perf build instead. Here it demonstrates the wrapper's failure mode: the common package is present, but the kernel-specific executable is not. A warning about a missing tool for the running kernel is a packaging problem, not a benchmark result. Install the matching kernel tools through your normal package-management process, then repeat these checks. Do not interpret a failed tool lookup as a slow workload.

Checkpoint

Continue only when perf --version identifies an executable that matches the kernel tools available on the host. The version printed above describes the local package, not a universal default.

2. Run the smallest scheduler smoke test

The scheduler pipe suite measures repeated pipe operations between two tasks. Run it with the documented default settings:

$ perf bench sched pipe
(executing 1000000 pipe operations between two tasks)

        Total time: 0.016 sec
                0.016000 usecs/op
                62500 ops/sec

Your numbers will differ substantially with CPU speed, kernel configuration, system load and the installed perf build. The useful shape is a statement of the operation count followed by total time, microseconds per operation and operations per second. A zero exit status tells you the command completed; it does not make the result comparable with a different workload.

The manual documents one million pipe operations as the default for this suite. The sample values above are illustrative output, not a promise about your machine. Record the whole output, kernel version, CPU topology and background load if you need a comparison that someone else can reproduce.

3. Make a short test for iteration

Use the suite's -l or --loop option to reduce the operation count while checking a command line:

$ perf bench sched pipe --loop 1000
(executing 1000 pipe operations between two tasks)

        Total time: 0.016 sec
                16.948000 usecs/op
                59004 ops/sec

This is useful for a smoke test, but it is a poor basis for a fine performance claim: fixed start-up and scheduling costs occupy more of a short run. Once the command is correct, choose a loop count that keeps the measurement long enough for your comparison and keep it unchanged across candidates.

Checkpoint

Verify that the operation count in the first line matches the value you selected. If it does not, stop and inspect the option spelling before comparing numbers.

4. Separate human output from script output

The common --format option accepts default for human-readable output and simple for a script-friendly result. A simple scheduler pipe run prints a single measurement:

$ perf bench --format=simple sched pipe
5.988

Do not mix the two forms in one data set. The simple value is intended for automated processing, while the default form gives context about operations and derived rates. Capture standard output and standard error separately if a script must reject warnings rather than store them as measurements.

5. Repeat a run without changing the workload

The common -r or --repeat option controls how many times a benchmark is repeated. Its documented default is 10:

$ perf bench --repeat=5 --format=simple sched pipe
5.988

Use an explicit repeat value in a comparison so a future default cannot silently change the amount of work. Check the installed command's output before writing a parser: the exact presentation of repeated results can be version-specific, while the option contract comes from the local manual.

Keep the host as quiet as practical. Do not run a benchmark during a deployment, backup or latency-sensitive incident unless you have accepted the interference. These suites create processes, threads or system calls for their test and can add measurable load even though they do not alter persistent configuration.

6. Compare scheduler modes carefully

The messaging suite exercises scheduler and inter-process communication paths. Its documented default uses 20 sender and receiver processes per group. You can switch to threads and choose the number of groups:

$ perf bench sched messaging --thread --group 20
(20 sender and receiver threads per group)
(20 groups == 800 threads run)

        Total time: 0.582 sec

Processes and threads are different workloads. The --pipe option selects pipe() instead of socketpair(); it does not mean that the whole suite becomes the pipe benchmark. Change one of --thread, --pipe or --group at a time, and write the selected values beside each result.

Do not begin with a large group count on a constrained host. The documented example reaches 800 threads, and a larger value can consume scheduler and memory capacity. Stop the command with the normal terminal interrupt if it is taking too long; an interrupted measurement is not a valid result, but it does not require an undo operation.

7. Inspect memory and system-call suites

Other subsystems cover different questions. The syscall basic suite measures throughput for a simple getppid(2) call. The mem memcpy suite measures copying, and mem memset measures setting memory. Start by listing the command's available benchmark help on a host with a working perf binary:

$ perf bench mem memcpy --help

For the memory suites, the manual documents a default size of 1 MB, architecture-dependent copy or set functions, and a loop count. The option name -l is used for a size in the documented memcpy and memset sections and for loops in the scheduler suites, so prefer the long option where the suite provides one and read the suite-specific help before scripting. The -c option selects perf CPU-cycle measurement instead of the default gettimeofday call for these memory tests.

Do not compare a CPU-cycle result directly with a wall-clock result. Note the measurement method, size, function and loop count in every record. Available functions depend on the architecture; the x86-64 names in the manual are not a portable list.

8. Treat cgroup tests as a separate, privileged decision

The scheduler pipe suite accepts --cgroups SENDER,RECEIVER to measure pipe operations between existing cgroups. Perf does not create or delete those cgroups, and the names must already exist and be accessible:

$ perf bench sched pipe --cgroups GROUP_A,GROUP_B

This example is ordinary command execution, but access to a particular cgroup hierarchy may require elevated privileges or a delegated cgroup controller. Check the paths and permissions first. Do not create or modify a production cgroup merely to obtain a benchmark point. If you do have an authorised test hierarchy, remove only the test cgroup through the same management system that created it, after confirming no process still depends on it.

Done means

  • The matching perf tool was checked against the running kernel and package version.
  • A scheduler pipe smoke test completed and its operation count was verified.
  • Repeat count, output format and workload options were explicit in recorded runs.
  • Process, thread, pipe and socketpair measurements were kept as separate workloads.
  • Memory results record size, function, loop count and measurement method.
  • No service, persistent configuration or production cgroup was changed for the test.