Test a SCSI Device Buffer Safely with sg_test_rwbuf

sg_test_rwbuf writes a pseudo-random pattern into a SCSI device's internal buffer, reads it straight back, and checks the two match. It never touches the medium itself, but a busy device can still make the test unsafe.

Budget 10 to 15 minutes: a read-only capacity check, one small write/read pass, and whatever time it takes to pin down the right device node.

This guide uses the installed sg_test_rwbuf from sg3-utils package version 1.46-3ubuntu4. The executable reports utility version 1.20 20191220 on this machine, while the installed manpage is marked sg3_utils-1.43. Treat the local command output as the authority for whichever options your own host has.

1. Identify the SCSI generic device

Find the device node before you run anything that can issue SCSI commands.

$ ls -l /dev/sg*
$ lsscsi -g

If nothing matches, stop here and resolve device discovery first.

Checkpoint: Write down the exact device path you intend to test. In the examples below, replace /dev/sg2 with that confirmed path.

2. Ask for the available buffer size

Use --quick to issue a READ BUFFER descriptor command and print the available buffer length and offset. It exits without running the write/read sequence:

$ sg_test_rwbuf --quick /dev/sg2
Read buffer size: 32768 bytes

The exact wording and numbers depend on the device; the result that matters is the available data-buffer length. If the command fails, record its diagnostic and exit status before changing anything:

$ sg_test_rwbuf --quick /dev/sg2
$ status=$?
$ printf 'sg_test_rwbuf status: %s\n' "$status"

A permission error normally means your account cannot open the device. Elevation may be required for the next command, but sudo does not identify the hardware or make a wrong device safe.

3. Check the safety boundary

The utility sends WRITE BUFFER in data mode and does not modify device microcode or the logical-unit medium. That does not make every situation harmless: the manpage warns that a device's buffer may already be in use, and concurrent activity can make data read back differently or cause corruption.

Warning: Do not test a disk whose mounted file systems are actively being used, a production path, or hardware carrying data you cannot recreate. Stop workloads and unmount file systems only within your normal change process. Unmounting changes system state, can interrupt services and is outside this guide; use your site's recovery procedure if you have already stopped a workload.

Before proceeding, confirm all three points:

4. Run one small write and read-back test

Choose a size below the reported capacity. A 4096-byte test is a conservative starting point when the buffer is at least that large:

$ sg_test_rwbuf --size=4096 --times=1 /dev/sg2
Success

The program writes a pseudo-random pattern to the internal buffer, reads it back and compares checksums over the requested size. A successful run reports Success and returns status 0. Verify the status immediately if this is part of a script:

$ status=$?
$ printf 'test status: %s\n' "$status"
test status: 0

The command normally needs access to the device node. If an unprivileged run reports permission denied, rerun the same confirmed command with sudo, subject to your site's privilege policy:

$ sudo sg_test_rwbuf --size=4096 --times=1 /dev/sg2

Do not add privilege merely because a command is a storage tool. Use the least privilege that can open the selected node.

5. Repeat only when the result is meaningful

--times repeats the write/read test; its default is 1. Repetition can help expose an intermittent adapter or buffer problem, but it also extends the period during which the device is under test:

$ sg_test_rwbuf --size=4096 --times=10 --verbose /dev/sg2

--verbose increases output. Keep the complete output and the device identity with the test record. A failed comparison may show up to 24 bytes from the first mismatch: one line is what was written and the next is what was received. That evidence is useful, but on its own it does not prove a faulty disk, cable or host adapter.

For an intermittent failure, first repeat the read-only --quick query and check that the device is still the expected one. Then inspect cabling, logs and workload activity using your normal hardware procedure. Do not keep retrying against a busy production device.

6. Use additional bytes only for a specific test

--addwr=AW writes AW additional zero bytes, while --addrd=AR reads AR additional bytes. The checksum still covers only the first --size bytes. These options are useful when reproducing a known adapter boundary issue, not as a routine default:

$ sg_test_rwbuf --size=4096 --addwr=16 --addrd=16 /dev/sg2

Make sure the total operation fits the device's buffer and record why the extra bytes are required. If you do not have a clear test case, omit both options.

7. Keep scripts explicit

For automation, check the exit status rather than parsing the word Success. Capture diagnostics and fail closed when the command returns non-zero:

if sg_test_rwbuf --size=4096 --times=1 /dev/sg2; then
    printf '%s\n' 'SCSI buffer test passed'
else
    status=$?
    printf 'SCSI buffer test failed with status %s\n' "$status" >&2
    exit "$status"
fi

Use a fixed, reviewed device path in automation. A changing /dev/sgN number can point at different hardware after a reboot or rescan; resolve stable identity as part of your deployment process rather than silently testing the first node returned by a glob.

Done means