Measure Linux CPU and Disk Activity with iostat
You will finish with repeatable CPU and block-device samples, a way to focus on one device, and enough context to tell a busy disk from a busy processor. The examples use iostat from sysstat 12.6.1-2, installed here as /usr/bin/iostat. Allow about 10 minutes for a first check, or longer if you need to compare a workload over time.
The route
Jump straight to the step you need, or tick off Done means at the end.
Before you start
You need a Linux shell with sysstat installed. Normal users can run every command in this guide; no elevated privileges are required. The command reads kernel statistics from /proc and /sys, so /proc must be mounted. You can confirm the binary and package version without changing the machine:
command -v iostat
iostat -V
Expected output includes /usr/bin/iostat and a line beginning with sysstat. The exact version and device names will differ between machines.
Checkpoint 1: take a baseline
Run iostat with no interval first:
iostat
This prints one CPU report and one device report. The first report covers activity since boot, not a fresh measurement of the command you just started. That history is useful for context, but it can hide a short-lived problem. The device list normally contains devices for which the kernel has statistics. To include every block device defined by the system, including unused devices, use ALL:
iostat ALL
On a wide report, the CPU row is an average across processors. The fields to scan first are %user, %system, %iowait, %steal and %idle. High %iowait says that CPUs were idle while an I/O request was outstanding; it does not by itself identify the slow device.
Checkpoint 2: sample a live workload
Use an interval and count when you need a bounded observation:
iostat -d -y 2 6
This prints six device reports, two seconds apart. The -d option selects device statistics and -y omits the since-boot report, so each displayed sample covers one two-second period. Without a count, the command continues until you stop it with Ctrl-C:
iostat -d -y 2
Stopping this display does not stop or alter the workload. It only ends iostat. For a quick CPU-only view, use iostat -c -y 1 5. If you omit -y, the first line can be a long historical average followed by short interval samples, which is an easy comparison error.
Checkpoint 3: narrow the device report
Pass device names after the options and before the interval. Replace the placeholders with names from the device column, such as sda, nvme0n1 or a device available on your host:
iostat -x -y DISK_DEVICE 1 5
For example:
iostat -x -y sda 1 5
The -x option adds extended metrics. The most useful starting points are r/s and w/s for completed reads and writes, rkB/s and wkB/s for throughput, await for average request service time including queue time, aqu-sz for average queue length, and %util for the proportion of elapsed time when requests were issued.
Interpret %util carefully. Near 100% can indicate saturation for a device that serves requests serially. It is not a universal performance limit for parallel devices such as modern SSDs or RAID arrays. Compare it with throughput, latency and queue length, and with the application behaviour you are investigating.
Make units and names explicit
By default, transfer rates use 1K blocks. In the iostat terminology, these are actually kibibytes, even though the headings say kB. Choose a unit deliberately when sharing a result:
iostat -d -k -y 1 3
iostat -d -m -y 1 3
iostat -d -h -y 1 3
-k selects kilobyte-style rates, -m selects megabyte-style rates, and -h uses human-readable values and the pretty layout. An inherited POSIXLY_CORRECT environment variable changes the default transfer unit to 512-byte blocks. Check it when two apparently similar reports disagree:
printf 'POSIXLY_CORRECT=%s\n' "${POSIXLY_CORRECT-}"
For stable device identity, request persistent names where the host provides them. This is useful when a device name such as sda can change after hardware or virtual-machine changes:
iostat -j ID -x -y 1 3
The persistent-name directories under /dev/disk must exist for the requested name type. With long persistent names, iostat automatically uses its pretty layout.
Export a machine-readable sample
Use JSON when another program will collect or compare the measurements:
iostat -o JSON -d -y 1 3 > iostat-sample.json
test -s iostat-sample.json && head -c 240 iostat-sample.json
The field order is not guaranteed and new fields can appear in later sysstat versions. Parse field names rather than relying on their order. Redirecting the report creates or replaces iostat-sample.json; choose a disposable directory or filename before running the command if an existing file matters. Nothing in iostat itself changes disk data, but the shell redirection can overwrite a file.
Common traps and safe recovery
- One report is not a live sample. It describes since-boot history. Use
-ywith an interval for workload-level observations. - A device report is not process attribution. iostat shows device totals. Pair it with a process-level tool when you need to find which program issued the I/O.
- Partitions can confuse totals. Use
-p DEVICEto show a device and its used partitions, for exampleiostat -p sda 2 3. Do not add partition and whole-device rates together without checking what each line represents. - Colour is not a diagnosis. Terminal colours are range indicators, not proof of an incident. Disable terminal colour when producing plain logs.
- A missing report often means missing kernel data. Check that
/procis mounted and that the named device exists. Do not fix an observation problem by changing storage configuration during an incident.
The examples are observational and do not require sudo. There is therefore no service rollback or configuration undo step. To remove the sample file created above, first verify its exact path and then use rm -- iostat-sample.json only if it is disposable. That deletion is irreversible through the shell.
Done means
- You can distinguish the since-boot report from interval samples.
- You can collect a bounded sample with
-y INTERVAL COUNT. - You can focus on a real device and interpret
await,aqu-szand%utiltogether. - You have checked units, persistent naming and JSON behaviour before comparing reports.