Trace syscalls and page faults with perf trace
By the end of this guide you will be able to run a bounded syscall trace, show only failed calls, inspect page faults and read a small summary of what a command did. The examples use the perf command from Debian's linux-tools-common package, version 6.8.0-142.142 was installed when this guide was written. Allow about 10 minutes if perf is already installed and usable.
The route
Jump straight to the step you need, or tick off Done means at the end.
Before you start
You need a shell, a command that you can run without changing important data, and a working perf binary. Check the package and command first:
dpkg-query -W -f='${Package} ${Version}\n' linux-tools-common
perf --version
The package version and the kernel-specific perf binary are separate concerns on Debian systems. If perf --version says that tools for the running kernel are missing, install the matching tools package through your normal system administration process before continuing. Do not treat the warning as a trace result.
Checkpoint
Continue when perf --version prints a version instead of a missing-tools warning.
1. Trace a short-lived command
Start with a command that exits by itself. The event stream normally includes system calls made by the command and, by default, inherited child activity. This example limits the output to four matching open-family calls:
perf trace -e 'open*' --max-events 4 true
The single quotes protect the wildcard from the shell. The selector is interpreted by perf trace, so it can match names such as openat. --max-events stops processing after the requested number of events. For strace-like output, the call line contains a process name and ID, the syscall name and arguments, its return value, and a duration in milliseconds. Exact lines depend on the libraries and filesystem used by your machine.
For a more useful test, trace a command that opens a known file:
perf trace -e 'open*' --max-events 8 sh -c 'cat /etc/hostname > /dev/null'
Checkpoint
You should see lines containing an openat-style call, followed by a return value. If the command produces no events, first check the selector with perf list and confirm that the kernel-specific tools are installed.
2. Narrow the trace to failures
Successful calls can bury the error you need. Add --failure and deliberately ask a harmless command to open a path that does not exist:
perf trace --failure -e 'open*' --max-events 4 sh -c 'cat /path/that/does/not/exist > /dev/null'
The command itself will print its own error, while perf trace reports only syscalls with a negative return value. This is useful when a service starts but cannot find a configuration file, socket or shared object. The trace does not repair the problem and it may expose pathnames, process names or other arguments, so save output only where that information is acceptable.
3. Include a syscall summary
When line-by-line output is too noisy, use --summary. It reports syscall timing by thread with minimum, maximum and average times, plus relative standard deviation:
perf trace --summary sh -c 'cat /etc/hostname > /dev/null'
Use --with-summary when you want both the individual calls and the summary. Add --errno-summary with either summary mode to include statistics for error numbers. These modes are often a better first pass than collecting a long unbounded trace.
4. Trace page faults separately
System calls are enabled by default. -F selects page faults, and its optional value is maj, min or all. Without a value, the installed manpage documents major faults as the default:
perf trace --no-syscalls -F maj --max-events 4 sleep 1
To include both minor and major faults:
perf trace -F all --max-events 8 sleep 1
A page-fault line identifies the fault type and may show the instruction symbol, the object containing the faulted address, whether the mapping is executable, and whether it belongs to the kernel. Symbol names are more useful when suitable debugging symbols are available. Do not read the displayed fault duration as handling time: the manpage currently documents that duration as always zero.
5. Observe an existing process
For a process that is already running, use its process ID with --pid. Find a test process in one terminal:
sleep 30
In another terminal, replace 12345 with the real process ID:
perf trace --pid 12345 --max-events 20
Stop tracing with Ctrl-C. This does not terminate the traced process, but attaching to a process you do not own may be denied by permissions or system security policy. Existing processes can also finish before the attach completes, so use a deliberately long-running test command while learning.
6. Add selected kernel events
The -e selector is not limited to syscalls. It can select tracepoints and other perf events. Ask perf list what this kernel exposes, then select a bounded set. For example, the manpage demonstrates limiting scheduler switches:
perf trace -e 'sched:*switch/nr=2/' --max-events 2
System-wide collection is a different scope. -a collects from all CPUs and may require elevated privileges. It can produce a large amount of data and affect a busy host, so begin with a short run and a narrow selector:
sudo perf trace -a -e 'sched:*switch/nr=2/' --max-events 2
Only use sudo when the requested scope requires it. Do not make broad system-wide tracing a default service setting. If you need to keep a trace for later processing, the installed manpage documents perf trace record as the shortcut that writes the raw syscall events needed for a perf data file; consult the matching perf record documentation for recording-specific options.
Common traps
- Unbounded output: use
--max-events, a narrow-eselector, or--summarybefore tracing a busy process. - Wrong event spelling: use
perf list; wildcard matching is supported, but available tracepoints vary by kernel and configuration. - Startup noise: use
--delay MILLISECONDSto wait after starting a program before measuring it. - Unexpected children: add
--no-inheritwhen child tasks should not inherit counters. - Missing symbols: install appropriate debugging symbols through your distribution, then repeat the trace. The trace can still be valid without them.
Done means
perf --versionworks with tools matching the running kernel.- You can bound a syscall trace with
-eand--max-events. - You can isolate failed calls with
--failureand aggregate timings with--summary. - You can distinguish syscall tracing from page-fault tracing with
--no-syscallsand-F. - You know that
--pidattaches to an existing process and-aexpands collection to every CPU.