Check Disk Health Safely with smartctl
You will identify a Linux disk, read its SMART health and error history, run a non-captive self-test, and check the result in about 10 minutes. The installed package here is smartmontools 7.4, released on 1 August 2023. The exact output varies by disk, protocol and USB or RAID controller.
The route
Jump straight to the step you need, or tick off Done means at the end.
Before you touch a disk
You need a shell, the smartmontools package and the device path. Reading information is usually harmless, but opening a device can require elevated privileges. A self-test is also a real operation: it can reduce performance while it runs and may take minutes or hours. Do not start one during a latency-sensitive workload without checking the device's estimate first.
These commands read data only:
smartctl --version
smartctl --scan
The first command should report smartctl 7.4 on this installation. The scan prints candidates such as /dev/sda -d scsi. It does not prove that you can read the device. If access is denied, repeat the later inspection commands with sudo.
Checkpoint
Write down the exact device path from the scan. Do not substitute a partition such as /dev/sda1 when the scan found the whole disk.
1. Read identity and health
Replace /dev/sdX below with the device you identified. This is a read-only inspection and normally needs root access on Linux:
sudo smartctl -i -H /dev/sdX
-i prints identity details including model, serial and firmware information. -H asks the device for its health status. A healthy result is useful evidence, not a guarantee that every future failure will be predicted. A failing health result means the drive may already have failed or may be predicting failure within roughly 24 hours. Stop treating it as a normal disk, copy data to a safe destination, and investigate the replacement path.
For NVMe, use the controller or namespace path reported by the scan, for example /dev/nvme0 or /dev/nvme0n1:
sudo smartctl -i -H /dev/nvme0
If autodetection is wrong, specify the type shown by the hardware path or controller documentation. For an ATA disk behind a SCSI-to-ATA translation layer, a common form is:
sudo smartctl -d sat -i -H /dev/sdX
Do not add -d sat at random. USB bridges and RAID controllers need different device types, and the wrong pass-through command can produce misleading errors. The local manpage documents the supported -d values for this 7.4 build.
2. Collect the useful evidence
For a first report, use -x. It requests SMART and non-SMART information, including capabilities, attributes and relevant logs. It is more useful than blindly relying on the older -a shortcut, which the manpage no longer recommends for ATA disks because it omits information that needs 48-bit ATA commands.
sudo smartctl -x /dev/sdX | tee smartctl-sdX.txt
The file records the report for later comparison. It contains identifiers such as the serial number, so treat it as operational data rather than posting it publicly without review.
To inspect the most useful parts separately:
sudo smartctl -A /dev/sdX
sudo smartctl -l error /dev/sdX
sudo smartctl -l selftest /dev/sdX
-A shows vendor-specific attributes. Do not compare raw values between manufacturers as though they shared one scale; the drive firmware defines their meaning. -l error shows the device error log, while -l selftest shows recorded self-test results. A log entry is a reason to investigate the disk, cabling, power and controller path, not by itself a diagnosis.
Checkpoint
Save the model, firmware, health result, attribute table, error log and self-test log before starting a test. A later report is easier to interpret when you can compare it with this baseline.
3. Check the test estimate
Ask the device which capabilities it advertises and how long its self-tests may take:
sudo smartctl -c /dev/sdX
Look for support and estimated duration for the short and extended tests. Some devices, bridges and virtualised environments expose only part of SMART. An absent feature is not evidence that the disk is failing.
4. Run a short self-test
Start with the short test. It checks electrical, mechanical and read performance and usually takes under ten minutes on ATA disks. The normal non-captive form can run during ordinary operation:
sudo smartctl -t short /dev/sdX
The command normally returns an estimated completion time. It starts the test rather than waiting for it. Avoid -C for routine checks: captive mode runs in the foreground and can make the device unavailable to normal I/O.
After the reported time has passed, read the result:
sudo smartctl -l selftest -H /dev/sdX
For a deeper check, schedule the extended test after checking -c:
sudo smartctl -t long /dev/sdX
A long test can take tens of minutes to several hours and may reduce performance. Do not start it simply because the short test passed. Schedule it for a maintenance window, then read -l selftest when it is complete. If you must stop a non-captive test, the documented abort command is:
sudo smartctl -X /dev/sdX
Power loss or shutdown should not damage the disk during a non-captive self-test, but the test may be aborted or resume when the device returns. The result may therefore be incomplete.
5. Use the exit status in scripts
Do not make automation parse a sentence such as SMART overall-health self-assessment test result. smartctl returns a bitmask. Zero means no reported problem. Non-zero bits distinguish command errors, access failures, failed health status, attributes at or below threshold, historical threshold failures, error-log records and self-test-log records.
Capture the status immediately after the command:
sudo smartctl -H /dev/sdX
smart_status=$?
if [ "$smart_status" -eq 0 ]; then
echo "SMART health check passed"
else
printf 'smartctl exit bitmask: %s\n' "$smart_status" >&2
fi
For a more specific check, bit 3 means the SMART status reported a failing disk, so its mask is 8:
if [ $((smart_status & 8)) -ne 0 ]; then
echo "SMART reports a failing disk" >&2
fi
Exit status 2 can also mean that the device could not be opened or identified, or that -n deliberately skipped a low-power device. Separate access and power-state failures from evidence of media failure in monitoring alerts.
Common traps
- Wrong path: inspect the whole device from
--scan, not a mounted partition. - USB or RAID masking: a bridge or controller may not pass through every SMART command. Check the controller's documented mapping and use the matching
-dmode. - Unexpected spin-up:
-n standbycan avoid checks for sleeping or standby ATA and SCSI devices. Autodetection itself may still wake a disk, so the manpage warns that an explicit device type may also be needed. - Raw attribute panic: names and raw values are vendor-specific. Compare the normalised value with its threshold and the manufacturer's guidance, then correlate it with logs and symptoms.
- Secret leakage: reports contain serial numbers and hardware identifiers. Redact them before sending a report outside the machine.
Done means
- You used
smartctl -i -Hagainst the correct whole-device path. - You recorded an evidence report with
-x, including attributes and logs. - You checked capabilities before starting a self-test.
- You read
-l selftestafter the test's estimated completion time. - Your script captures and interprets the exit bitmask rather than treating every non-zero result as media failure.