Home / Alt manpages / perf-sched(1)

  • perf-sched(1)
  • User command
  • linux

Measure Scheduler Latency with perf sched

You will record a short workload, inspect the scheduler events it produced, and narrow the report to a CPU, process or time window when the trace is too large. The useful measurements are wait time, scheduling delay and run time, all shown in milliseconds with microsecond precision.

Allow about fifteen minutes. You need a Linux shell and a matching perf executable. Recording scheduler events may need elevated privileges, depending on the kernel's perf permissions. The examples do not change services or persistent configuration, but a recording can contain process names and timing information, so treat the resulting perf.data as sensitive operational data.

Version checkpoint: the local package is linux-tools-common 6.8.0-142.142. On this machine, /usr/bin/perf reports that the matching tool for the running kernel, 6.8.0-139, is missing. If you see the same warning, install or select the matching kernel-tools package through your normal change process before expecting a recording to work. Do not diagnose a missing binary as a scheduler problem.

1. Check the installed command

Start with read-only checks. They do not need sudo:

$ command -v perf
/usr/bin/perf
$ perf --version
WARNING: perf not found for kernel 6.8.0-139
$ dpkg-query -W -f='${Package} ${Version}\n' linux-tools-common
linux-tools-common 6.8.0-142.142

Your warning and package version may differ. The important check is that the perf program matches the running kernel closely enough for your distribution. If perf --version prints a version rather than a warning, keep that output with your investigation notes.

Checkpoint: run perf sched or read man 1 perf-sched. The subcommands covered here are record, latency, map, replay, script and timehist.

2. Record a small, repeatable workload

Use a short command first. This records scheduler events while sleep runs:

$ perf sched record -- sleep 1

The double hyphen makes the boundary between perf sched options and the workload obvious. The command is normally unprivileged on permissive systems. If the kernel denies access to performance events, rerun the same command with sudo only after checking your local perf policy:

$ sudo perf sched record -- sleep 1

By default, later reports read perf.data. Check that it exists before moving on:

$ test -s perf.data && echo 'recording is non-empty'
recording is non-empty

A recording is a file, not a live dashboard. Do not overwrite an existing trace that you may need. Move it to a deliberate name first, or choose a separate working directory:

$ mv perf.data perf-data-scheduler-test.data
$ perf sched timehist -i perf-data-scheduler-test.data

This move is reversible with mv perf-data-scheduler-test.data perf.data. Avoid deleting a trace until its results and retention requirements are clear.

3. Read the event-level report

With the default file name, the shortest report command is:

$ perf sched timehist
            time    cpu  task name             wait time  sch delay   run time
                         [tid/pid]                (msec)     (msec)     (msec)
    ...

The installed manual describes three useful columns. Wait time is the interval between a task being scheduled out and its next scheduled-in event. Scheduling delay is the interval between wakeup and the task actually running. Run time is how long it ran for that event. Times are printed as seconds, a dot and microseconds in the event timestamp, while these duration columns are in milliseconds.

Do not treat one large value as a system-wide average. Idle tasks, kernel threads and short-lived workload threads all appear in the event stream. The task name is accompanied by a thread or process identifier, which is the safer thing to use when filtering.

4. Produce a compact summary

When the event list is too noisy, ask for a summary by thread:

$ perf sched timehist --summary
      ... summary by thread with min, max, average and relative stddev ...

The exact rows depend on the trace. The summary reports minimum, maximum and average run times in seconds, plus relative standard deviation. Use --with-summary when you need both the event list and the per-thread summary:

$ perf sched timehist --with-summary

These options analyse the existing recording. They do not record a second workload and do not change scheduling policy.

5. Narrow the trace without recording again

Use filters when a trace includes several CPUs or processes. For example, this reports events for CPUs 0 and 3:

$ perf sched timehist --cpu 0,3

For a known process or thread, use its numeric identifier:

$ perf sched timehist --pid 12345
$ perf sched timehist --tid 12345

Use --time for a timestamp window taken from the report. Both endpoints use seconds and microseconds; either end may be omitted:

$ perf sched timehist --time 79371.870000,79371.900000
$ perf sched timehist --time ,79371.900000
$ perf sched timehist --time 79371.870000,

The first command includes only the chosen interval. The other forms start at the beginning or continue to the end of the file. Copy timestamps accurately: a typo can produce an empty report that looks like a workload failure.

6. Inspect the trace in other views

perf sched map prints a text outline of context switches. Columns represent CPUs, two-letter shortcuts represent tasks, an asterisk marks the CPU where an event occurred, and a dot marks an idle CPU:

$ perf sched map
$ perf sched map --compact

--compact is useful on high-core-count hosts because it shows only CPUs with activity. Use --cpus when you need to focus the map on selected CPUs. The underlying map remains text, so it is still usable in logs and terminals without visual highlighting.

For a detailed event trace, use:

$ perf sched script

The manual describes this as an alias of perf script for now. Use it when the summary tells you that a delay exists but you need the event sequence around it.

7. Handle failure and clean up safely

If the command says that perf is not found for the running kernel, check the running kernel and the distribution's matching tools package. Installing a package is an administrative change, so use your normal approval and maintenance process:

$ uname -r
6.8.0-139-generic
$ apt-cache policy linux-tools-$(uname -r)
  Candidate: ...

If the candidate is unavailable, do not guess a package name or mix tools from an unrelated kernel. Ask the distribution administrator to provide the matching package, then repeat step 1. If recording is denied, inspect the kernel's perf permission policy and the account used for the test before adding sudo.

When you are finished, remove only the trace you have identified as disposable. This is irreversible:

$ rm -- perf-data-scheduler-test.data

Keep the file instead if another person needs to reproduce the report, and restrict its access according to the process names and timing information it contains.

Done means

  • The installed perf matches the running kernel, or the mismatch is recorded as the blocker.
  • A short workload produced a non-empty scheduler trace.
  • perf sched timehist showed wait time, scheduling delay and run time.
  • You used a summary or a CPU, PID, TID or time filter to reduce noise where needed.
  • You kept, moved or removed perf.data deliberately rather than overwriting useful evidence.