Measure Kernel Work Latency with perf kwork
You will record a short, harmless workload and use perf kwork to inspect kernel work: how long work ran, how long it waited, and which work categories or CPUs were involved. The examples use the perf-kwork(1) interface supplied by linux-tools-common version 6.8.0-142.142 on this machine.
The route
Jump straight to the step you need, or tick off Done means at the end.
Allow about fifteen minutes. You need a shell, a readable perf installation matching the running kernel, and enough permission for perf events. The workflow writes a perf.data trace in the current directory. It does not change kernel settings, services or persistent configuration. Tracing can expose workload and system details, so keep the trace file private.
1. Check the installed tool before recording
Start with read-only checks. They do not need elevated privileges:
$ command -v perf
/usr/bin/perf
$ dpkg-query -W -f='${Package} ${Version}\n' linux-tools-common
linux-tools-common 6.8.0-142.142
$ perf --version
The package and command can be present while the version-specific perf binary is missing for the running kernel. On this host the wrapper reports that perf is not installed for kernel 6.8.0-139. That is a prerequisite failure, not a reason to guess at the output. Install the matching distribution package through your normal change process, then rerun the checks. Do not use sudo just to read the version.
Checkpoint: continue only when perf --version returns a version instead of a missing-tools warning.
2. Record a bounded workload
Change to a directory where you can keep the trace, then record one second of an idle sleep:
$ mkdir -p "$HOME/perf-kwork-test"
$ cd "$HOME/perf-kwork-test"
$ perf kwork record -- sleep 1
[ perf record: Woken up 1 times to write data ]
[ perf record: Captured and wrote ... bytes perf.data (... samples) ]
The two final lines are representative perf progress messages; the byte count and sample count vary. The -- makes it clear that sleep 1 is the workload passed to the recorder. The recorder writes perf.data, so check that it exists before analysing:
$ test -s perf.data && echo 'perf.data is ready'
perf.data is ready
This command is normally unprivileged. If perf reports a permissions error, inspect the host's perf policy and the command's diagnostic first. Raising privileges may broaden access to system-wide activity and sensitive data; use it only under an approved troubleshooting procedure.
3. Inspect runtime and delay
Run the default report against the file just created:
$ perf kwork report
Runtime start Runtime end Cpu Kwork name Runtime Delaytime
(TYPE)NAME:NUM (msec) (msec)
... individual kernel-work events ...
The report is event-oriented. Runtime is the time spent executing the work. Delaytime is the interval between the work being raised and its actual entry. Times are displayed in milliseconds with microsecond precision. Your event names, CPU numbers and rows will differ because they describe this kernel and this run.
For a compact view grouped by useful measurements, select a sort key and ask for the summary:
$ perf kwork report --sort runtime --with-summary
$ perf kwork report --sort max --with-summary
$ perf kwork report --sort count --with-summary
The report supports runtime, max and count sort keys. Use one command at a time while learning the output; changing several filters at once makes an empty result harder to diagnose.
4. Measure latency rather than execution time
Use the latency subcommand when waiting time is the question:
$ perf kwork latency
$ perf kwork latency --sort avg --with-summary
Latency has its own sort keys: avg, max and count. Do not read a high runtime as a high scheduling delay, or a high delay as proof that the work itself is expensive. Compare both reports when deciding whether the pressure is execution time or queueing.
To restrict analysis to a named kernel-work item, add --name NAME. The name must match an item present in your trace. For example, first inspect the unfiltered report, then try:
$ perf kwork latency --name SCHED
If this produces no rows, remove the filter and copy the name from the report. Do not infer names from a process name or from a guess about the kernel version.
5. Narrow by CPU or time window
Use a comma-separated CPU list when one processor matters:
$ perf kwork report --cpu 0,1
$ perf kwork latency --cpu 3
The time filter uses seconds and microseconds. A missing start or stop means the beginning or end of the file:
$ perf kwork report --time 1.000000,2.000000
$ perf kwork latency --time ,1.500000
$ perf kwork latency --time 2.000000,
These values refer to timestamps in the recorded data, not wall-clock times on your watch. A filter outside the trace naturally returns no useful events. If that happens, rerun without --time and use the timestamps shown by the report to choose a window.
6. Compare with timehist and top
timehist analyses kernel-work events in a timeline and can show call chains when they are available:
$ perf kwork timehist
$ perf kwork timehist --call-graph --max-stack 8
--max-stack limits the displayed backtrace depth; the documented default is five functions. Symbols may require a matching vmlinux, kallsyms file or symbol directory, selected with --vmlinux, --kallsyms or --symfs. Missing symbols reduce detail but do not turn a latency report into a runtime report.
Use top for task CPU usage rather than per-work latency:
$ perf kwork top
$ perf kwork top --sort runtime
The top report has its own sort keys: rate, runtime and tid. Keep this distinction in your notes when comparing results.
7. Finish safely and clean up deliberately
When the investigation is over, remove the trace only if it contains no evidence you need. This is irreversible:
$ rm -- "$HOME/perf-kwork-test/perf.data"
If you need to retain it, restrict access instead:
$ chmod 600 "$HOME/perf-kwork-test/perf.data"
$ ls -l "$HOME/perf-kwork-test/perf.data"
The report commands do not modify the trace. Re-run them with different filters while perf.data remains available.
Done means
- The perf binary matches the running kernel and returns a real version.
- A bounded workload produced a private, non-empty
perf.data. - You inspected runtime and delay separately with
reportandlatency. - CPU and time filters were chosen from the trace rather than guessed.
timehistandtopwere used for timeline and task-usage questions.- The trace was either protected for later analysis or deliberately removed.