Capture a Focused Kernel Trace with trace-cmd record

Resist the urge to trace everything: trace-cmd record rewards a narrow, deliberate capture far more than a broad one. You will record a short Ftrace capture into trace.dat, restrict it to a useful event, and read the result with trace-cmd report.

The examples use trace-cmd 3.2.0 from Ubuntu package 3.2-1ubuntu2. Allow about fifteen minutes, including a little time to identify an event that exists on your kernel.

Tracing reads and changes kernel tracing state. Use a test machine or a quiet maintenance window, keep captures short, and do not start with every event. Most systems require elevated privileges to access the tracing filesystem. The commands below use sudo for recording only where that access is needed.

1. Confirm the command and an event

Check the installed version, then ask trace-cmd which events are available. This is a read-only check and does not start tracing:

$ trace-cmd --version
trace-cmd version 3.2.0 (not-a-git-repo)
$ trace-cmd list -e sched_switch
sched:sched_switch

Your event list can differ. The -e argument accepts an event such as sched_switch, a subsystem such as sched, or the form subsystem:event-name. The special value all enables every event, but that can create a large, noisy capture and is a poor first diagnostic.

If the list command says permission is denied while reading /sys/kernel/tracing/available_events, do not guess that the event is absent. Check the mount and retry the read-only query with privilege:

$ findmnt /sys/kernel/tracing
$ sudo trace-cmd list -e sched_switch

Checkpoint: you have an event name printed by the local command, or you have recorded that tracing access is unavailable before attempting a capture.

2. Record a short command

Run a short-lived workload and write the capture to an explicit path. The double hyphen separates trace-cmd options from the command being traced:

$ sudo trace-cmd record -e sched_switch -o /tmp/sched-switch.dat -- sh -c 'sleep 1'
  trace-cmd record: Waking up 4 threads
  trace-cmd record: Latency: ...
  trace-cmd record: saved 1.0 MB in ...

The exact progress lines and size vary. On success, verify the file and its type:

$ test -s /tmp/sched-switch.dat && echo 'capture exists and is non-empty'
capture exists and is non-empty
$ trace-cmd report /tmp/sched-switch.dat | head -12

When a command is supplied, recording ends when that command exits. Without a command, trace-cmd record continues until you press Ctrl-C. That interactive form is useful for reproducing an issue, but set an explicit stop time or watch the output file size for longer runs.

Warning: by default, trace-cmd resets the buffers and disables the tracing it enabled when recording ends. Do not add -k casually: it keeps the tracer and buffers from being reset, and the manual warns that it leaves tracing_on set to zero. If you used -k during debugging, plan a separate cleanup with the appropriate trace-cmd reset operation and verify the host's tracing state afterwards.

3. Read the capture

Use trace-cmd report to turn the binary capture into readable event lines:

$ trace-cmd report /tmp/sched-switch.dat | head -8
           swapper/0-0     [000] ...: sched_switch: prev_comm=...
             sleep-1234    [001] ...: sched_switch: prev_comm=...

Field formatting depends on the kernel and the recorded activity. A successful report proves that the file is structurally readable, not that the trace contains the event you expected. Search the report for the event name and the workload:

$ trace-cmd report /tmp/sched-switch.dat | grep -E 'sched_switch|sleep' | head

Tip: if you need wall-clock correlation, add --date to the recording command. The option writes timestamps into the trace buffer after recording so the timestamps can be mapped to gettimeofday when the file is read.

4. Narrow the capture when it is too noisy

Choose a subsystem or event rather than enabling everything. For example, this records all events in the scheduler subsystem:

$ sudo trace-cmd record -e sched -o /tmp/sched.dat -- sh -c 'sleep 1'

You can exclude matching events after selecting a subsystem. The exclusion applies to events specified after -v:

$ sudo trace-cmd record -e sched -v -e '*stat*' -o /tmp/sched-no-stat.dat -- sh -c 'sleep 1'

For a function trace, use -p function or -p function_graph and limit it with one or more -l function names. Avoid --func-stack with an unfiltered function tracer: the manual warns that collecting a stack for every function without a successful function filter can live-lock the machine.

Event filters belong immediately after their event. For example, the following asks the kernel to retain scheduler switches whose previous priority is greater than 100:

$ sudo trace-cmd record -e sched_switch -f 'prev_prio > 100' -o /tmp/high-prio-switches.dat -- sh -c 'sleep 1'

Whether a field and expression are accepted depends on the kernel event definition. If the command rejects the filter, inspect the event format or remove the filter; do not silently treat an unfiltered capture as equivalent.

5. Control size and system impact

Check the host after the run and stop if the tracing workload affects the service under investigation.

6. Diagnose a failed or incomplete recording

A permission error at tracing_on, available_events or another file points to tracing filesystem access, not necessarily a bad event name. Inspect the mount and retry with sudo if that is permitted by your operating policy. If the event itself is missing, trace-cmd normally exits with an error; -i tells it to ignore named events that are not found, which is useful only for deliberately portable scripts where missing data is acceptable.

If the output file is absent or empty, check the command's exit status and stderr before interpreting the report:

$ sudo trace-cmd record -e sched_switch -o /tmp/check.dat -- sh -c 'sleep 1'
$ printf 'record exit status: %s\n' "$?"
record exit status: 0
$ stat --format='%n %s bytes' /tmp/check.dat

Warning: do not overwrite a useful capture while experimenting. Use a new filename, and remove old files only after you have copied the evidence you need. A trace can contain sensitive command names, process IDs, kernel addresses or timing information, so handle the resulting .dat file according to the same policy as logs from the host.

Done means