Safely Check and Control SCSI Reset Escalation with sg_reset

sg_reset can check whether SCSI reset recovery is already underway, or trigger one of four increasingly disruptive resets yourself. The guide uses the installed sg_reset from sg3-utils 1.46-3ubuntu4. The binary reports version 0.67 (20200501), while the local manual page is labelled sg3_utils 1.43, so check the command on your own host before relying on older examples.

Allow about fifteen minutes for a read-only check, plus a maintenance window for any real reset. You need a Linux shell, an appropriate SCSI device node such as /dev/sg2, and permission to open it. A reset can interrupt I/O, affect other logical units, or re-initialise the host adapter, so do not experiment against a production path just to see what happens.

1. Confirm the installed command

Start with the two read-only queries below. They do not contact a SCSI device and do not need elevated privileges:

$ command -v sg_reset
/usr/bin/sg_reset
$ sg_reset --version
sg_reset: version: 0.67 20200501
$ sg_reset --help
Usage: sg_reset [--bus] [--device] [--help] [--host] [--no-esc] [--no-escalate] [--target]
                [--verbose] [--version] DEVICE

Your help output may wrap differently. The important detail is that a device argument is always required, even for a status check. The long options are easier to audit in a runbook, though the short forms exist too.

Checkpoint: Record the path printed by command -v and the version shown by --version. If this is not the expected installation, stop before choosing a reset.

2. Check reset state without changing it

Run sg_reset with only the device node. With no reset option, it simply asks whether reset recovery is underway:

$ DEVICE=/dev/sg2
$ sg_reset "$DEVICE"
$ status=$?
$ printf 'reset-state check exit status: %s\n' "$status"
reset-state check exit status: 0

A successful exit status means the check completed. It is not a promise the device is healthy, mounted, or free of application-level errors. A non-zero status means the check failed or the device reported a problem: keep the diagnostic text and investigate the path before sending a reset.

For a deliberate harmless probe on a non-SCSI file, the installed command fails without issuing a reset at all:

$ sg_reset /dev/null
sg_reset: SG_SCSI_RESET failed: Inappropriate ioctl for device
$ printf 'exit status: %s\n' "$?"
exit status: 1

This is a test of error handling only. Replace /dev/null with the actual SCSI node before interpreting the result for real.

3. Identify the reset scope before acting

The options describe increasingly broad scopes along the path to DEVICE:

Start at the smallest scope that addresses the fault. Do not choose --host because it is easy to remember, and do not assume a bus reset touches only the named device: the low-level driver and transport decide what the request actually does.

4. Prevent automatic escalation

Without --no-esc, a failed device reset can escalate to a target reset, then a bus reset, then a host reset. That default can cause collateral disruption across a large storage configuration. Pair the requested scope with --no-esc when you need exactly one attempt:

$ sudo sg_reset --device --no-esc "$DEVICE"
sudo password: ********
$ printf 'reset request exit status: %s\n' "$?"
reset request exit status: 0

--no-escalate is the equivalent long spelling, and the installed help also shows the short form -N. Elevated privilege is only an example here: use it when the device permissions require it, and follow your host's change-control process. Do not add sudo merely to make a failed check look successful.

Warning: This command changes live recovery state and may interrupt I/O. There is no undo command for a reset. Stop affected services first when your storage design requires it, confirm multipath or failover behaviour, and make sure you can still reach the host through an independent management path.

5. Use broader scopes only with an explicit reason

If the device-level recovery path is known to be stuck and your storage documentation calls for a target reset, make the wider scope visible in the command:

$ sudo sg_reset --target --no-esc "$DEVICE"
$ printf 'target reset exit status: %s\n' "$?"
target reset exit status: 0

Use the same pattern for --bus or --host only after assessing the other devices sharing that path:

$ sudo sg_reset --bus --no-esc "$DEVICE"
$ sudo sg_reset --host --no-esc "$DEVICE"

These are operational commands, not a ladder to try automatically. A zero status reports that the request was accepted by the relevant interface; it does not prove every affected device recovered. Recheck the reset state, then verify the storage stack using the tools appropriate to your multipath, filesystem or service configuration.

6. Interpret failures and keep recovery bounded

A failure such as Inappropriate ioctl for device usually means the path is not a supported SCSI device node for this operation. Check the node with your normal inventory tools and confirm it is the intended device. Permission errors call for an access review, not a wider reset.

The manual page notes that support depends on the Linux SCSI mid-layer and low-level driver. Modern transports may treat a bus reset as a dummy operation, or implement it as resets of related targets. A successful command therefore describes the request path, not a universal physical action. Transport-specific link resets may need a different tool supplied by the HBA or enclosure vendor.

After a reset, keep the original command, timestamp, exit status and diagnostic output. Check the device again without a reset option:

$ sg_reset "$DEVICE"
$ printf 'post-reset check exit status: %s\n' "$?"
post-reset check exit status: 0

If I/O remains unhealthy, stop repeating resets. Escalate to the storage owner with the affected path, shared devices, kernel messages and the exact scope requested. Repeated host or bus resets can turn a single-device fault into a much wider outage.

Done means