Find CPU Hotspots Live with perf top
You will finish with a live profile showing which symbols are consuming CPU, plus a small set of filters for turning that broad view into a useful diagnosis. The examples follow the installed perf-top(1) manual from linux-tools-common version 6.8.0-142.142. Allow about fifteen minutes for a first investigation.
The route
Jump straight to the step you need, or tick off Done means at the end.
You need a shell, a matching perf executable, and a workload that is currently doing something worth measuring. Most examples are ordinary commands. Reading system-wide counters or kernel symbols may require elevated privileges or a suitable perf_event_paranoid policy. Do not change that policy or grant capabilities just to make one experiment work without first checking your system's security requirements.
1. Check the installed tool before profiling
Start with read-only checks. They identify the command and the kernel-tools mismatch that is a common source of wasted time:
$ command -v perf
$ dpkg-query -W -f='${Package} ${Version}\n' linux-tools-common
$ uname -r
$ perf --version
On this machine, linux-tools-common is 6.8.0-142.142, while the installed wrapper reports that it cannot find a kernel-specific perf for kernel 6.8.0-139. If you see that warning, stop here and install or otherwise provide the matching package through your normal operating-system process. This guide does not install packages.
Checkpoint
Continue only when perf --version identifies a usable binary and does not report a missing kernel-specific tool.
2. Confirm the event names available on this host
Event names are hardware and kernel dependent. List the names instead of guessing them:
$ perf list
Look for a hardware CPU event such as cycles. The manual also accepts raw PMU events in the form rN, but raw encodings depend on the PMU format exposed under /sys/bus/event_source/devices/cpu/format/. Use a symbolic event first. If the list is empty or opening an event fails, treat that as an access or hardware-support problem, not as a reason to invent an event name.
3. Start a short system-wide view
With no process selector, perf top uses system-wide collection by default. That is useful for finding the busiest code on a test host, but it can expose activity from other users and services:
$ sudo perf top -e cycles
The sudo is conditional. Try the same command without it first on a non-sensitive machine. If the kernel denies access, read the error and ask an administrator about the existing perf policy. System-wide profiling is observational, but it can reveal process names, paths, kernel addresses and workload behaviour. Treat the display as sensitive operational data.
The terminal opens an interactive display. You should see a changing table of symbols, shared objects and overhead percentages. The exact rows depend on the workload. Press q or Q to quit. This does not create a file and does not change the workload.
Checkpoint
You have a live table and can exit it cleanly before adding more options.
4. Narrow the profile to one process
System-wide output is often too noisy. Find the target process ID, then select it with -p:
$ pgrep -x YOUR_PROGRAM
$ sudo perf top -p YOUR_PID -e cycles
Replace YOUR_PROGRAM and YOUR_PID with values from your host. The process must still exist when perf top opens it. The manual also accepts a comma-separated list of process IDs. If the program creates short-lived workers, process-wide results may change as those workers appear and disappear; record the command, PID and time when comparing runs.
To inspect only selected CPUs instead, use a comma-separated list or range:
$ sudo perf top -C 0,1 -e cycles
Do not confuse -C, which selects CPUs, with -p, which selects processes. Both examples still refresh an interactive display.
5. Make the display easier to read
Refresh interval and row count are safe presentation controls. For example, refresh every two seconds and show 20 functions:
$ sudo perf top -e cycles -d 2 -E 20
Press d to change the delay interactively, e to change the number of entries, and f to set a hit-count display filter. Use --percent-limit when you want to suppress entries below a percentage, but remember that hiding rows does not reduce the measured activity.
For a focused view, filter by symbols, shared objects or commands with --symbols, --dsos or --comms. Filtering changes how the overhead percentage is interpreted. The manual calls the two modes relative and absolute: relative percentages are recalculated from the filtered entries, while absolute percentages retain their original share. Check this before comparing a filtered screen with an unfiltered one.
6. Add call graphs only when they answer a question
Use -g or --call-graph when the hot symbol alone is not enough and you need its callers. A simple starting point is:
$ sudo perf top -p YOUR_PID -e cycles -g
With callchains, the default Children view accumulates the cost of child functions into their parents. Its percentages can therefore add to more than 100 percent. Self is the direct sampled cost at the symbol. Press s to annotate a symbol, and S to return to the full profile. Annotation needs symbols and, for kernel code, may need a suitable vmlinux path supplied with -k.
Call-graph collection and symbol resolution increase overhead and may fail under tighter permissions. Start without -g, identify a real hotspot, then add it for a short confirmation run. Use --no-children if you need the traditional self-overhead ordering.
7. Diagnose a result before changing anything
A high percentage is a lead, not a fix. Repeat the observation with the same workload, event, process selector and delay. A row labelled unknown, an absent kernel symbol or an event-open error can indicate missing symbols, insufficient privileges, a mismatched perf binary or unsupported hardware.
Use verbose diagnostics when an event will not open:
$ sudo perf top -v -p YOUR_PID -e cycles
Do not make the output public if it contains process names, executable paths or kernel details. Stop with q; there is no recorded data to remove unless you separately started another command that writes a file. Do not use --force merely to silence an ownership warning: the option disables ownership validation and should have a specific, understood purpose.
Done means
perf --versionidentifies a usable tool matched to the running kernel.- You selected an event from
perf listinstead of guessing. - You can distinguish system-wide, process and CPU selection.
- You checked whether permissions, symbols or tool mismatch explain a failure.
- You can read the difference between
SelfandChildrenoverhead. - You quit the interactive session and made no persistent system change.