Measure Linux lock contention with perf lock
You will record a bounded workload, inspect the resulting lock statistics, and narrow the report to the locks or threads that need attention. Allow 15 to 30 minutes for a first run, plus the time needed to reproduce the workload. This guide covers the perf lock command documented with Debian's linux-tools-common package.
The route
Jump straight to the step you need, or tick off Done means at the end.
Before you start
You need a Linux host with perf and its matching kernel tools installed, a workload that is safe to run, and enough permission to collect the required perf events. The installed package on this machine is linux-tools-common 6.8.0-142.142. Its wrapper currently reports that no perf binary is installed for kernel 6.8.0-139, so the command examples below are checked against the installed perf-lock(1) manual but cannot be executed on this host until matching tools are installed.
Check your own installation before planning a measurement:
$ command -v perf
$ perf --version
$ man perf-lock
If perf --version prints a kernel-tools warning or the command is missing, stop there and install the package that matches the running kernel through your normal package-management process. Do not work around that warning by copying an unrelated binary into a system directory.
Perf access is also governed by the kernel's perf event security settings. A collection that fails as an ordinary user may need a permitted capability or an administrative invocation, depending on the host policy. Treat captured traces as potentially sensitive: they can expose process names, addresses and call paths. Use a directory with suitable permissions and remove the data when it is no longer needed.
1. Record a small, reproducible workload
Run the workload through perf lock record. The command creates perf.data in the current directory and records lock events from the start of the child command to its exit:
$ mkdir -p "$HOME/perf-lock-run"
$ cd "$HOME/perf-lock-run"
$ perf lock record -- /path/to/workload --one-test-case
$ test -s perf.data && echo "recording created"
recording created
Replace the placeholder command with a real, bounded workload. Quote paths and arguments rather than assembling an option string from untrusted input. The -- marker makes the boundary between perf lock options and the child command clear; the documented form is perf lock record <command>.
This example writes a new file. It does not change the workload or kernel configuration, but an existing perf.data may be replaced. Preserve an earlier capture first if it matters:
$ test ! -e perf.data || cp --preserve=all perf.data perf.data.previous
Recording the whole system is a different boundary. If your installed command supports it for the contention workflow, perf lock contention --use-bpf --all-cpus collects from all CPUs and can observe unrelated users and services. Use that only with operational approval and an explicit output directory. A process-specific measurement is easier to interpret and exposes less data.
Checkpoint: confirm the capture
Before analysing, verify that the file exists and that the recording command returned successfully:
$ printf 'record exit status: %s\n' "$?"
record exit status: 0
$ ls -lh perf.data
The size is workload-dependent. A zero-length or absent file means the collection did not complete; fix that first instead of treating an empty report as evidence that the workload has no lock activity.
2. Read the aggregate report
Use perf lock report to read statistical data from the default input file, perf.data:
$ perf lock report
Name acquired contended avg wait (ns)
mutex-name ... ... ...
Exact names and numbers depend on the workload. The report's useful questions are how often a lock was acquired, how often acquisition had to wait, and how long those waits took. Sort by another documented key when the default order hides the problem:
$ perf lock report --key=contended
$ perf lock report --key=wait_total
$ perf lock report --key=wait_max
The available report keys are acquired, contended, avg_wait, wait_total, wait_max and wait_min. A large acquisition count alone is not contention. Start with contended or a wait-time key, then compare the result with the workload's latency.
3. Change the report fields and scope
Keep output small when you are comparing runs. The --field option accepts the same report fields and can be repeated as a comma-separated value:
$ perf lock report --field=contended,avg_wait,wait_max
$ perf lock report --threads --field=acquired,contended,avg_wait
--threads changes the view to per-thread lock statistics. Use it when a process-wide total tells you that contention exists but not which worker is involved. --entries=20 limits the displayed number of entries, which is useful for an initial pass:
$ perf lock report --threads --entries=20 --key=contended
Use --combine-locks when several instances should be merged by lock class name. This makes a class-level hotspot easier to see, but it removes instance-level detail. Repeat the report without that option before deciding which address or call site to change.
4. Inspect raw events and metadata
When the aggregate is not enough, ask for raw events or metadata from the same capture:
$ perf lock script
$ perf lock info --threads
$ perf lock info --map
script shows raw lock events. info --threads lists threads in the data, while info --map shows the address-to-name map for lock instances. These outputs are more verbose and can contain identifiers that should not be sent outside the host without review.
To analyse a different capture, set the input explicitly:
$ perf lock report --input=/path/to/perf.data
$ perf lock script --input=/path/to/perf.data
Do not assume that a report is reading the file you intended. The default is perf.data, unless standard input is a FIFO. An explicit path is safer in scripts and when several captures are in the same directory.
5. Use live BPF contention collection only when needed
The contention subcommand can collect statistics with BPF instead of reading an existing capture:
$ sudo perf lock contention --use-bpf --entries=20
$ sudo perf lock contention --use-bpf --threads --key=contended
These commands may require elevated privileges and can observe activity beyond your shell. Use sudo only when the host policy requires it, and obtain approval before collecting on a shared production system. The BPF map has a documented default maximum of 16384 entries, with --map-nr-entries available when a measured workload needs a different bound. --max-stack defaults to 8 and --stack-skip defaults to 3; changing them affects collection cost and attribution, so record any non-default values with your results.
Filters reduce the result set. For example, --type-filter=mutex selects a lock type, --lock-filter=NAME selects addresses or names, and --callstack-filter=STRING matches a substring in the call stack. A call-stack filter of irq is deliberately broad: it can match more than one symbol. Test a filter against an unfiltered run before drawing a conclusion from an empty result.
6. Clean up and recover from common mistakes
There is no undo operation for an analysis because the normal commands read the capture. Remove only captures you have finished reviewing, and keep a copy if the result supports an incident or performance investigation:
$ rm -- perf.data
That deletion is irreversible unless another copy exists. If you kept the earlier backup from step 1, restore it with:
$ mv -- perf.data.previous perf.data
If a report complains about a missing input, use --input or return to the directory containing perf.data. If collection is denied, check the host's perf event policy and memory limits rather than repeatedly adding sudo. If the report is empty or unhelpfully broad, shorten the workload, choose a process-specific recording, or use a narrower filter. A high wait maximum from one outlier is a reason to inspect the raw events, not proof that every request is slow.
Done means
- Matching perf tools are installed and the version or kernel warning is understood.
- A bounded workload produced a non-empty
perf.datacapture. - You checked contention and wait time, rather than acquisition count alone.
- You used per-thread, metadata or raw-event views to narrow an interesting result.
- Any system-wide or BPF collection was authorised, and captured data is stored with appropriate permissions.
- You know which capture files to retain and which can be removed safely.