A flaky disk is hard to diagnose from silence, and scsi_logging_level turns up the kernel's SCSI chatter just enough to see it, reversibly. The examples use scsi_logging_level from sg3-utils version 1.46-3ubuntu4 on this machine; the installed command itself reports its script version as 1.0.
Allow about ten minutes. You need a shell and the sg3-utils package. Reading and previewing are ordinary user operations. Applying a setting changes kernel-wide SCSI diagnostic output, so save sudo for the one step that actually writes it.
Start with the read-only operation. It fetches the values from the kernel and prints the pseudo-file they came from:
$ scsi_logging_level --get
Current scsi logging level:
/proc/sys/dev/scsi/logging_level = 0
SCSI_LOG_ERROR=0
SCSI_LOG_TIMEOUT=0
SCSI_LOG_SCAN=0
SCSI_LOG_MLQUEUE=0
SCSI_LOG_MLCOMPLETE=0
SCSI_LOG_LLQUEUE=0
SCSI_LOG_LLCOMPLETE=0
SCSI_LOG_HLQUEUE=0
SCSI_LOG_HLCOMPLETE=0
SCSI_LOG_IOCTL=0
Your own values may differ. Each field accepts an integer from 0 to 7: zero is quiet, higher values produce more diagnostic output. This is system-wide for anything using the Linux SCSI subsystem, not a per-disk switch.
Checkpoint: save this output somewhere tied to the incident before changing anything. scsi_logging_level --get > /tmp/scsi-logging-before.txt records the baseline for this boot. Treat that file as temporary diagnostic data, and remove it after the investigation if it holds anything you would rather not keep around.
Use --create to parse the options and display the resulting fields without touching the kernel. It is the safety check to run before anything elevated:
$ scsi_logging_level --create --error=5 --timeout=0
/proc/sys/dev/scsi/logging_level = 5
SCSI_LOG_ERROR=5
SCSI_LOG_TIMEOUT=0
SCSI_LOG_SCAN=0
SCSI_LOG_MLQUEUE=0
SCSI_LOG_MLCOMPLETE=0
SCSI_LOG_LLQUEUE=0
SCSI_LOG_LLCOMPLETE=0
SCSI_LOG_HLQUEUE=0
SCSI_LOG_HLCOMPLETE=0
SCSI_LOG_IOCTL=0
That first line reflects the value that would be written, not a write that has already happened. Options can name individual fields, or grouped pairs: --midlevel sets both mid-level queue and completion fields, --lowlevel sets both low-level fields, --highlevel sets both high-level fields. The more specific option always wins when the same field is named twice: --all=1 --hlqueue=3 leaves everything else at 1 and sets only SCSI_LOG_HLQUEUE to 3.
Do not mistake --create for a dry-run wrapper around arbitrary commands: it only prepares and displays SCSI logging fields. Exactly one of --create, --get or --set is required per run.
Once the preview looks right, apply the smallest useful change. This example raises SCSI error logging while explicitly keeping timeout logging off:
$ sudo scsi_logging_level --set --error=5 --timeout=0
[sudo] password for your-user:
$ scsi_logging_level --get
Current scsi logging level:
/proc/sys/dev/scsi/logging_level = 5
SCSI_LOG_ERROR=5
SCSI_LOG_TIMEOUT=0
The abbreviated form is documented too: sudo scsi_logging_level -s -E 5 -T 0. Keep the full option names in runbooks where clarity matters more than typing speed. The write needs superuser permission and only changes the live kernel setting: there is no persistent configuration file, and nothing here survives a reboot.
Use a level no higher than the evidence actually requires. Excessive logging fills system logs and buries the one message you actually wanted. The command itself never restarts a device or service, though SCSI and driver messages can get noisy while a high level is active.
The sg driver reads the timeout field for its own logging. For a short, active investigation, push it to 7:
$ sudo scsi_logging_level --set --timeout=7
$ scsi_logging_level --get | grep '^SCSI_LOG_TIMEOUT='
SCSI_LOG_TIMEOUT=7
Watch the system log with your host's normal log reader while reproducing the problem. Upstream sg3_utils documentation points at the system log, often /var/log/syslog, as the destination, though exact messages and locations depend on kernel, driver and distribution.
When the test ends, turn the extra trace off immediately:
$ sudo scsi_logging_level --set --timeout=0
$ scsi_logging_level --get | grep '^SCSI_LOG_TIMEOUT='
SCSI_LOG_TIMEOUT=0
Recovery: if timeout logging was already non-zero in your baseline, restore that recorded value rather than assuming zero. The same rule applies to every field: recovery is just another --set using the values you captured in step 1.
Levels outside the inclusive range 0 to 7 are rejected before anything is applied:
$ scsi_logging_level --set --error=8
scsi_logging_level: log level '8' out of range, expect '0' to '7'
scsi_logging_level: Try 'scsi_logging_level --help' for more information.
A permission error means rerunning the same command with sudo, but check the options with --create first. If the command cannot reach /proc/sys/dev/scsi/logging_level, check that the proc filesystem is mounted and that the running kernel actually exposes the SCSI logging interface. Do not create or edit that pseudo-file blindly: its availability is a property of the kernel and the proc mount, not something you can fix by hand.
In scripts, check the exit status rather than matching displayed text. Zero means success; anything else is an error. Keep read, preview, apply and restore as separate steps so a failed diagnostic never gets mistaken for a successful change.
sg3-utils version and current SCSI fields are recorded.--create before elevation.