Home / Alt manpages / perf-c2c(1)

  • perf-c2c(1)
  • User command
  • linux

Find False Sharing with perf c2c

You will record a short workload run, then inspect which cachelines attract the most cache-to-cache traffic and HITM accesses. That gives you a starting point for finding false sharing or other shared-data contention. Allow 15 to 30 minutes, including a second run if the first capture is too noisy.

This guide uses perf c2c from the installed linux-tools-common package, version 6.8.0-142.142, and the perf-c2c(1) manual dated 1 September 2026. The actual events are hardware and kernel dependent. The installed /usr/bin/perf currently reports that the tool for kernel 6.8.0-139 is missing, so the checks below also show how to recognise that packaging problem before you spend time debugging a workload.

Recording can expose process names, addresses, source locations and timing data. Use a test workload whose data you are allowed to profile, and do not copy perf.data into a place where unrelated users can read it.

1. Check the matching perf installation

Start with ordinary, unprivileged checks. You need the perf command, a workload that runs long enough to sample, and a CPU with a supported performance-monitoring facility. The command may need elevated privileges, depending on the kernel's perf policy and the workload you select.

$ dpkg-query -W -f='${Package} ${Version}\n' linux-tools-common
linux-tools-common 6.8.0-142.142
$ perf --version
WARNING: perf not found for kernel 6.8.0-139

On a usable installation, perf --version prints a version instead. The warning means the common package is present but the kernel-specific perf binary is not. Install the matching linux-tools-6.8.0-139-generic package through your normal system-management process, then repeat the check. Do not work around a kernel mismatch by copying a random perf binary from another machine.

2. Record one controlled workload

Run the workload as the final argument to perf c2c record. The following example records a harmless three-second sleep, which verifies the command path but is not expected to produce useful application contention:

$ perf c2c record sleep 3

For a real investigation, replace the final command with a bounded test case. Keep the capture short and repeatable:

$ perf c2c record -- /path/to/test-workload --seconds 10

The -- separates c2c options from options passed to the ordinary perf record command. The c2c subcommand configures extra sampling options itself, including physical data and sample CPU information. Without an -e option, the manual documents Intel load and store events, AMD IBS, and PowerPC load and store events. Arm64 uses SPE and requires suitable hardware and kernel support.

A successful capture normally leaves perf.data in the current directory. A run can fail with a counter-open error even when the syntax is correct: the CPU may not provide the event, the kernel may restrict perf, or the matching tool may be absent. Use -v when you need those errors displayed.

3. Check the capture before interpreting it

Confirm that the file exists and belongs to the run you just made:

$ stat -c '%n %s bytes' perf.data
perf.data 123456 bytes
$ test -s perf.data && echo 'capture is non-empty'
capture is non-empty

The byte count is workload dependent, so the number above is illustrative. A missing or empty file is a recording failure, not evidence that the program had no contention. Keep the file until the report has been checked. If you deliberately want a different output location, use the ordinary perf record output option after the separator, and confirm the resulting path before starting analysis.

4. Produce a script-friendly report

The report command opens a TUI by default. Use --stdio for a terminal log, a remote shell or a saved investigation note:

$ perf c2c report --stdio

The report contains overall memory-access statistics, a shared cacheline table and a distribution of offsets within those cachelines. The first table is the useful triage list. Look for high HITM counts or percentages, then inspect the offsets and the process, instruction, shared object and source line responsible for the accesses.

If the capture is not in the current directory, name it explicitly:

$ perf c2c report --stdio --input /path/to/perf.data

The default display is total HITM, except that Arm64 defaults to peer mode. Use -d rmt, -d lcl or -d peer when you need a particular view. --stats restricts the output to statistic tables. These choices change how you read the percentages; they do not create extra samples.

5. Narrow the noisy parts

When the report identifies a cacheline, coalesce its offsets to make the relevant dimension easier to see:

$ perf c2c report --stdio --coalesce pid
$ perf c2c report --stdio --coalesce iaddr
$ perf c2c report --stdio --coalesce dso

The documented fields are tid, pid, iaddr and dso; the default is pid,iaddr. Use --show-all if a small percentage is relevant. Otherwise the report omits captured HITM lines below its documented 0.0005 percent threshold, which can hide a real but infrequent event.

For a system-wide capture, pass -a to the underlying perf record command. That requires care because unrelated processes will be included:

$ perf c2c record -- -g -a sleep 10

This example also records callchains. It can require more permission and storage, and it makes attribution noisier. Prefer a process-scoped workload first. If the report needs kernel symbols, pass a matching uncompressed kernel image with -k /path/to/vmlinux. Do not guess the image: mismatched symbols produce confident-looking but wrong source locations.

6. Treat HITM as a lead, not a verdict

HITM means that a load observed a modified copy in another cache, so it is a useful signal for cacheline contention. It does not by itself prove false sharing. Compare the offset, owning data structure, thread behaviour and workload phases. A cacheline can be legitimately shared, and sampling means the counts are not a complete audit of every memory operation. Arm SPE is explicitly statistical in the manual, and other architectures also depend on their PMU event semantics.

For a false-sharing hypothesis, map the reported offset back to the structure layout, then change one suspected field's placement or access pattern in a test branch and repeat the same bounded capture. Keep the original perf.data so the before and after runs are comparable. Do not change production alignment, scheduling or kernel settings solely because one short report has a high percentage.

Common failure traps

  • Missing matching tool: a kernel-specific warning from perf is a packaging issue. Install the package matching the running kernel and verify the version again.
  • Permission denied or counter-open errors: check the system's perf policy and event support. Do not immediately use sudo on a sensitive workload; elevated recording can expose more kernel and process information.
  • No useful HITM rows: lengthen the bounded workload, reduce unrelated activity, or inspect the overall statistics. An empty result does not prove that no sharing exists.
  • Overwritten evidence: recording normally writes perf.data. Move or rename a capture before a new run if it matters. Removing it is irreversible, so retain it until the report and any comparison are complete.

Done means

  • The perf binary matches the running kernel and the required PMU support is available.
  • A short, repeatable workload produced a non-empty perf.data.
  • perf c2c report --stdio identified cachelines, offsets and responsible code where samples were available.
  • HITM percentages were treated as sampled evidence, then checked against structure layout and workload behaviour.
  • Captures were kept private and were not overwritten before the investigation was finished.