Home / Alt manpages / sg_write_long(8)

  • sg_write_long(8)
  • Admin command
  • linux

Test SCSI Medium Errors Safely with sg_write_long

You will finish with a controlled way to send SCSI WRITE LONG to one logical block, verify the resulting read error, and restore the block. The command can write logical data together with ECC bytes, or mark a block uncorrectable without transferring data. This is a storage fault-injection tool, not a normal file-writing command.

Allow about thirty minutes, plus time for a real backup and for the device to report its long-block size. You need the sg3-utils package, a SCSI device that accepts WRITE LONG, and root access for the final device operations. The installed package here is sg3-utils 1.46-3ubuntu4; the binary reports sg_write_long 1.21 20180723. The installed manpage is labelled sg3_utils-1.42, so trust the local binary's help for its exact option display and the manpage for the documented command contract.

Warning

Every write example below changes a physical device. A wrong device, LBA or transfer length can damage data or a filesystem. Do not use a mounted production disk. Substitute a disposable test device for /dev/sdX, and stop if you cannot identify it unambiguously.

1. Confirm the installed command

This first step is read-only and does not need elevated privileges:

$ command -v sg_write_long
/usr/bin/sg_write_long
$ sg_write_long --version
sg_write_long: version: 1.21 20180723
$ sg_write_long --help
Usage: sg_write_long [--16] [--cor_dis] [--help] [--in=IF] [--lba=LBA]
                     [--pblock] [--verbose] [--version] [--wr_uncor]
                     [--xfer_len=BTL] DEVICE

The command sends either WRITE LONG(10), the default, or WRITE LONG(16) with --16. WRITE LONG(10) only carries a 32-bit LBA. Use --16 when the target LBA cannot fit in 32 bits.

Checkpoint: continue only when the version and option names match the command you are about to run. Do not copy examples from a different sg3-utils release without checking its help.

2. Choose the test block and long length

Set two placeholders in your notes. TEST_LBA is the logical block to test, and BTL is the device's long-block size, including ECC. The documented default is 520 bytes, but that is not a safe assumption for every device. A mismatch causes the command to write nothing and report an error response that can help identify the required length.

DEVICE=/dev/sdX
TEST_LBA=0x1234
BTL=520

Use an exact whole-device path, not a partition, for a test that you have designed around a physical SCSI block. Check it before proceeding:

# lsblk -o NAME,TYPE,SIZE,FSTYPE,MOUNTPOINTS /dev/sdX
# lsscsi -g

These inspection commands are ordinary reads. The # prompt marks commands that are commonly run as root; use sudo if your shell is not already privileged. Confirm that no filesystem or service uses the target device.

3. Save the original long block

Do this before sending WRITE LONG. The companion sg_read_long utility saves the logical data and ECC bytes, giving you a direct restoration file:

# sg_read_long --lba=0x1234 --out=0x1234-original.img --xfer_len=520 /dev/sdX

Replace both the LBA and length with your recorded values. The local manpage notes that you may need more than one read attempt to discover the correct length. If this step fails, stop. Do not guess a length and continue with a write.

Checkpoint: keep the image in a separate, protected location and verify that it exists:

$ stat --printf='%n %s bytes\n' 0x1234-original.img
0x1234-original.img 520 bytes

4. Send a deliberately bad long block

There are two different tests. This one transfers BTL bytes of 0xff by default, including the long-block data, and may result in a medium error when the device later reads the block:

# sg_write_long --lba=0x1234 --xfer_len=520 /dev/sdX

The command normally returns no output on success and exits with status 0. Capture that status immediately:

$ printf 'sg_write_long status: %s\n' "$?"
sg_write_long status: 0

Alternatively, on a device that supports it, --wr_uncor flags the LBA as uncorrectable and transfers no data:

# sg_write_long --lba=0x1234 --wr_uncor /dev/sdX

That option is not a harmless simulation. It changes the device's error state and support varies, so use it only when the test plan specifically needs an uncorrectable mark.

5. Read the result through the SCSI pass-through path

Use sg_dd with blk_sgio=1 to read exactly one block through SG_IO rather than letting the block layer hide the device error:

# sg_dd if=/dev/sdX blk_sgio=1 skip=0x1234 of=/tmp/sg-write-long-check.bin bs=512 count=1 verbose=4

An error response or non-zero exit is the expected result for a successful fault-injection test. Exact sense text depends on the device. If the read succeeds, that does not prove the write was ignored: some devices reject the malformed long data, recover it, or expose different error behaviour. Preserve the verbose output with the test record.

6. Restore the saved block

Restoration is another device write. Write the saved long image back with the same LBA and long length:

# sg_write_long --lba=0x1234 --in=0x1234-original.img --xfer_len=520 /dev/sdX

A successful command normally prints nothing and returns 0. Check the block again with the same sg_dd command. If you only need a known-good block and the device's logical block size is 512, the manpage also documents writing one zeroed logical block with sg_dd, but that discards the original contents:

# sg_dd if=/dev/zero of=/dev/sdX blk_sgio=1 seek=0x1234 bs=512 count=1

Use that alternative only when overwriting the block with zeros is intentional. There is no undo for it unless you have a separate backup.

7. Avoid automatic sector reassignment during recovery

If your test changes only a few bits and the device can correct the data, its read-write error-recovery settings matter. The manpage warns that ARRE or AWRE can cause a recovered error to reassign the LBA and add the old location to the grown defect list. That consumes a finite spare sector and is not easily reversed.

Inspect those settings before a recoverable-error experiment. Clearing them is itself a device configuration change and can last until power cycling:

# sdparm -c AWRE,ARRE /dev/sdX

Do not run that command casually. Review the device documentation and test plan first, and record the original settings if you change them. The --cor_dis option can inhibit correction and related recovery mechanisms for devices that support it, but it does not make an unsafe target safe.

Done means

  • The installed version and target device were recorded.
  • The original long block was saved before the write.
  • The LBA, command form and transfer length were checked.
  • The read test used blk_sgio=1 and its output was retained.
  • The original image was written back and the block was read again.
  • Any ARRE or AWRE change was deliberate, documented and reviewed.