Home / Alt manpages / cache_writeback(8)

  • cache_writeback(8)
  • Admin command
  • linux

Recover Dirty Device-Mapper Cache Blocks with cache_writeback

You will prepare an offline device-mapper cache recovery and write dirty blocks from the fast device back to the origin device. The installed command is cache_writeback 0.9.0 from thin-provisioning-tools 0.9.0-2ubuntu5.1. Allow at least fifteen minutes for the checks, then as long as the cache contents and devices require.

This is a recovery operation, not a routine cache-management command. The manual describes it for cases where the SSD is giving I/O errors. You need the cache metadata, the slow origin device, and the fast device containing cached data. You must also be able to access those paths, normally as root. Do not run this against a live cache.

1. Confirm the installed command

Start with read-only checks. They do not open your devices for writeback:

$ command -v cache_writeback
/usr/sbin/cache_writeback
$ cache_writeback --version
0.9.0
$ cache_writeback --help
Usage: cache_writeback [options]
        --metadata-device <dev>
        --origin-device <dev>
        --fast-device <dev>
        --buffer-size-meg <size>
        --list-failed-blocks
        --no-metadata-update

The help output above is from the installed binary. The option spelling is worth checking on the recovery host rather than copying a command intended for a different thin-provisioning-tools release.

Checkpoint

Stop here if the command is missing, the version is not the one you have assessed, or the recovery plan does not identify all three inputs.

2. Keep the cache offline

Stop every workload that could use the cached mapping, and make sure the cache is not presented as a live device-mapper cache. Unmount filesystems and stop dependent services according to your normal incident procedure. This command cannot be run on a live cache, and writing while another process is using the mapping can make the recovery result unsafe.

There is no enable, disable, or dry-run option in this tool. The first invocation that names the devices is intended to perform writeback. Treat the origin and fast-device arguments as destinations and sources, not as interchangeable labels.

Record the exact paths before using elevated privileges:

$ META='/path/to/cache-metadata'
$ ORIGIN='/dev/mapper/origin-volume'
$ FAST='/dev/mapper/fast-cache-volume'
$ printf 'metadata: %s\norigin: %s\nfast: %s\n' "$META" "$ORIGIN" "$FAST"
$ ls -l -- "$META" "$ORIGIN" "$FAST"

Replace every placeholder with a path from your recovery record. Check that the metadata path identifies cache metadata, the origin path identifies the slow device being cached, and the fast path identifies the device that contains the dirty cached data. A syntactically valid path can still refer to the wrong volume.

3. Choose how metadata should be handled

The normal operation writes dirty blocks to the origin and updates metadata to clear the dirty flags. That is the setting to use when you are restoring the origin as the authoritative copy and intend to finish the cache recovery.

--no-metadata-update suppresses the metadata update that clears those flags. The manual gives decommissioning the cache as a possible reason to use it. This flag does not make the device operation read-only, and it does not make a failed write safe to ignore. Use it only when your recovery plan explicitly requires the metadata to remain unchanged.

Warning

Do not add --no-metadata-update because you are uncertain. Pause and confirm the intended metadata state with the person responsible for the cache layout. If the cache is being decommissioned, preserve the metadata and device evidence until the replacement or recovery has been checked.

4. Run the writeback once

With the cache offline and the paths independently verified, run the operation. It requires elevated access when the device nodes or metadata file are not accessible to your user:

# cache_writeback \
    --metadata-device "$META" \
    --origin-device "$ORIGIN" \
    --fast-device "$FAST"
# status=$?
# printf 'cache_writeback exit status: %s\n' "$status"

Do not paste this until the variables contain the real paths. A successful process exit is the first checkpoint, but it is not a substitute for checking the origin and any failed-block report. The tool is intentionally offline, so a service outage is expected for the duration of the recovery.

The command has no documented progress format in the installed manual. Avoid assuming that a quiet terminal means that no work is happening. Wait for the process to exit, and monitor it only with methods approved for your incident procedure.

5. List blocks that did not write back

If the writeback reports failures, or if your procedure requires an explicit failed-block listing, run the same operation with --list-failed-blocks as part of the command:

# cache_writeback \
    --metadata-device "$META" \
    --origin-device "$ORIGIN" \
    --fast-device "$FAST" \
    --list-failed-blocks

The option asks the tool to list blocks that failed the writeback process. The exact list is host- and incident-specific, so do not invent an expected block count. Save its output with the incident record and stop before bringing services back if any blocks failed.

Do not erase, recreate, reformat, or repurpose the fast device to make the error disappear. Those actions can destroy the remaining source data needed for recovery. Preserve both devices until the failed blocks have been assessed and the recovered origin has been verified.

6. Use the buffer option only when justified

--buffer-size-meg sets the size of the metadata cache used by the command. The installed manual gives a default of 16 Gig and says that a larger size may improve performance. It does not provide a universal value for a particular device size or memory budget.

Leave the default in place unless your recovery plan has a measured reason to change it. If you do change it, record the value in the incident notes and keep the rest of the command identical:

# cache_writeback \
    --metadata-device "$META" \
    --origin-device "$ORIGIN" \
    --fast-device "$FAST" \
    --buffer-size-meg 32768

The value above is an example of the option's shape, not a recommendation for your host. Check available memory and the tool's behaviour under controlled conditions before choosing a larger buffer during an active incident.

7. Bring the service back only after verification

When the command exits without reported failed blocks, verify the recovered origin using checks appropriate to the filesystem or application. A generic check is to confirm that the expected device mapping and mount are present, then use the application's own integrity check. Those checks are outside cache_writeback, so their output depends on your storage stack.

Only after that verification should you recreate or re-enable the cache, if your recovery plan still calls for one, and remount or restart the dependent service. There is no undo command that reverses blocks already written to the origin. If the operation was interrupted or failed, keep the cache offline, preserve the command output, and follow the storage team's recovery procedure instead of repeatedly guessing at device arguments.

If you used --no-metadata-update, do not treat a successful writeback as permission to discard the metadata. The dirty flags were deliberately left alone. Resolve that metadata state as a separate, documented decommissioning or repair step.

Done means

  • The installed cache_writeback version and all three paths were recorded.
  • The cache was offline for the entire writeback operation.
  • Metadata, origin and fast-device roles were checked independently.
  • The exit status and any failed-block listing were saved.
  • The origin was checked before services or a replacement cache were brought back.
  • No device was erased or repurposed while recovery evidence was still needed.