Home / Alt manpages / perf-script-python(1)

  • perf-script-python(1)
  • User command
  • linux

Build a Python Trace Script with perf script

You will record a focused Linux trace, generate a Python handler skeleton from the recorded events, and turn it into a small aggregation script. Allow 20 to 30 minutes for a first run. You need the perf tools, a writable working directory, and enough privilege to record the events you choose. This guide follows the installed perf-script-python(1) manual from linux-tools-common version 6.8.0-142.142.

The examples use system-wide syscall entry tracing. That can produce a large file quickly, so use a short capture and check the output before extending it. Recording is normally the only step that may need elevated privileges.

1. Check the installed perf tool

Start in an empty directory so that the capture and generated script are easy to identify:

$ mkdir -p "$HOME/perf-python-example"
$ cd "$HOME/perf-python-example"
$ perf --version

You should get a version line. On this machine, /usr/bin/perf is present but reports that the kernel-specific linux-tools-6.8.0-139-generic package is missing, so the command cannot be run successfully until the matching tool is installed through the normal package-management process. Do not confuse the manual package version with a working perf binary for the running kernel.

Checkpoint: continue only when perf --version exits successfully. If the command is missing or names a missing kernel-tools package, stop here and install the matching package according to your distribution's procedure.

2. Record a short event stream

Record syscall-entry events for every CPU for a few seconds:

$ sudo perf record -a -e raw_syscalls:sys_enter -- sleep 5
[ perf record: Captured and wrote ... perf.data (... samples) ]

The -a option makes collection system-wide. The event is one raw tracepoint for syscall entry, so the event's id field distinguishes calls. The manual's longer interactive example records until you press Ctrl-C; the bounded sleep command is easier to review and limits the capture.

This creates perf.data in the current directory. It contains activity from all users and processes observed during the capture. Treat it as potentially sensitive: process names, IDs, timing and other trace fields may be present. Keep it local and remove it when it is no longer needed.

Checkpoint:

$ test -s perf.data && printf '%s\n' 'perf.data is ready'
perf.data is ready

3. Generate the starter script

Generate handlers from the event types in the capture:

$ perf script -g python
generated Python script: perf-script.py
$ test -s perf-script.py && printf '%s\n' 'starter script generated'
starter script generated

The file is a diagnostic skeleton. It adds the perf support-module path, imports the helper modules, and creates a function for each event type found in perf.data. Event handlers are named subsystem__event_name. For this capture, the relevant handler is named raw_syscalls__sys_enter.

The installed manual's generated example uses Python 2-style print statements. Do not blindly add Python 3 syntax or assume that an arbitrary system Python is the interpreter used by perf. First inspect the generated file and test it against the installed perf build.

$ sed -n '1,120p' perf-script.py
$ perf script -s perf-script.py | sed -n '1,12p'

Expected output is a series of event lines containing the event name, CPU, timestamp, process ID, command name, and event fields. The exact values depend on the capture. If the output is empty, check that perf.data is in the current directory and that the recording actually collected samples.

4. Replace printing with an aggregation

Copy the starter before editing it, then add a counter and print totals at the end. The manual documents autodict() for this purpose and the syscall_name() helper for turning raw IDs into names:

$ cp --preserve=all perf-script.py syscall-counts.py
syscalls = autodict()

def trace_end():
    for syscall_id, count in sorted(syscalls.iteritems(),
                                    key=lambda item: (item[1], item[0]),
                                    reverse=True):
        print "%s %d" % (syscall_name(syscall_id), count)

def raw_syscalls__sys_enter(event_name, context, common_cpu,
                            common_secs, common_nsecs, common_pid,
                            common_comm, id, args):
    try:
        syscalls[id] += 1
    except TypeError:
        syscalls[id] = 1

Keep the imports and support-path setup from the generated file. Insert the counter at module level and replace the generated event handler with the version above. trace_end() runs after all events have been processed, so the report is emitted once rather than once per syscall.

Run the script against the existing capture:

$ perf script -s syscall-counts.py | sed -n '1,12p'
read 923
close 3037
write 455067

The counts and order will differ. The useful check is that the command exits successfully and prints syscall names followed by integer counts. If the handler is never called, regenerate the skeleton from the current perf.data; a script generated for different event types can report unhandled events or no useful totals.

5. Handle failures without losing the capture

Keep perf.data until the script has produced a report you trust. A Python error does not require a new recording. Fix the script and rerun:

$ perf script -s syscall-counts.py > syscall-counts.txt
$ status=$?
$ printf 'perf script exit status: %s\n' "$status"
$ test "$status" -eq 0

When you are finished, remove the trace only after checking that the report is useful:

$ rm -- perf.data perf-script.py syscall-counts.py syscall-counts.txt

This deletion is irreversible. If the capture may be needed for another analysis, archive it with an access-controlled method instead. Neither the recording nor the report should be copied to a shared location without checking the data they contain.

6. Keep a reusable script separate from recording

For a longer-lived tool, record only the tracepoints that the script needs. Use perf list to find candidate events and inspect the corresponding files under /sys/kernel/tracing/events/ for their fields. Generate a fresh skeleton after changing the event set, then keep the Python script under version control with a short note naming the required events.

The manual also describes installing paired -record and -report shell scripts under the perf source tree's scripts/python/bin directory. That is an installation and packaging workflow, not a requirement for running a local script. Do not modify a system perf installation merely to make a private script appear in perf script -l.

Done means

  • perf --version succeeds for the running kernel's matching tools.
  • A short, deliberately scoped recording produced perf.data.
  • perf script -g python generated handlers from that exact capture.
  • The event handler aggregates data and trace_end() prints the report.
  • You checked the report before deleting or sharing the potentially sensitive trace.