Repair Device-Mapper Cache Metadata Safely with cache_repair
You will finish with a repaired copy of device-mapper cache metadata written to a file or metadata device, while keeping the source untouched. The installed command is cache_repair from thin-provisioning-tools 0.9.0-2ubuntu5.1, and the binary reports version 0.9.0. Allow about fifteen minutes for the command itself and the checks around it. The repair may take longer on large metadata, and deciding whether the source is safe to use is an administrator's responsibility.
The route
Jump straight to the step you need, or tick off Done means at the end.
This guide is for offline binary metadata. It does not repair a mounted or active cache. You need a shell, read access to the input, write access to the destination, enough free space, and a maintenance plan if the destination is a real metadata device.
1. Confirm the installed command
Start with read-only checks. These do not need elevated privileges unless your system hides the binary or its version information from your account:
command -v cache_repair
cache_repair --version
cache_repair --help
On this installation, the useful version output is:
0.9.0
The help text shows the required shape: an input selected with -i or --input, and an output selected with -o or --output. Each value can name a device or a file. Do not confuse this program with cache_check, which is a separate tool for checking metadata.
2. Stop before touching live metadata
Checkpoint
Establish that the metadata source is not live before running the repair. The manual explicitly says that cache_repair cannot run on live metadata. Stop the cache using the procedure for your system, confirm that the relevant device-mapper target is no longer using the metadata, and only then continue. The exact stop command depends on how the cache was created and managed, so do not invent one from this guide.
This boundary matters because cache_repair reads binary metadata and writes a repaired result elsewhere. A successful exit status is not permission to run it against an active metadata device. If you cannot prove that the source is offline, do not run the repair.
3. Inspect the input and choose a new destination
Record the exact source and destination before adding sudo. In the example below, replace the obvious placeholders with paths from your maintenance plan:
INPUT=/path/to/offline-cache-metadata
OUTPUT=/path/to/repaired-cache-metadata
printf 'input: %s\noutput: %s\n' "$INPUT" "$OUTPUT"
test -r "$INPUT" && printf '%s\n' 'input is readable'
test ! -e "$OUTPUT" && printf '%s\n' 'output does not already exist'
These checks do not prove that the input contains valid cache metadata. They catch two common distractions: a mistyped path and an output name that would overwrite an existing file. Keep the original input. Do not use the same path for both variables.
If the output is a regular file, the manual requires it to be preallocated and large enough to hold the metadata. Allocate it using your normal storage procedure, based on the capacity you have deliberately chosen. If the output is a logical volume or other metadata device, verify its identity and capacity separately before writing to it. A wrong device path is a destructive error, not a repair failure that can be undone by rerunning the command.
4. Run the repair into the separate destination
Once the source is offline and the destination has been checked, run the ordinary command:
cache_repair --input "$INPUT" --output "$OUTPUT"
status=$?
printf 'cache_repair exit status: %s\n' "$status"
For a successful run, the status is 0. The tool's normal diagnostic contract is simple: it returns 0 for success and 1 for an error. It is not safe to interpret a missing or truncated output as repaired merely because a shell command was entered. Check the captured status immediately, before running another command.
Use elevated privileges only when the source or destination permissions require them:
sudo cache_repair --input "$INPUT" --output "$OUTPUT"
status=$?
printf 'cache_repair exit status: %s\n' "$status"
Warning
sudo changes who can open the paths; it does not make live metadata safe, increase a destination's capacity, or select the correct device. Read the command again before pressing Enter, especially the value after --output.
5. Verify the result before putting it back into service
First check the destination using the storage tools appropriate to its type. For a regular file, confirm that it exists and has a non-zero size:
test -s "$OUTPUT" && stat --printf='repaired output: %n (%s bytes)\n' "$OUTPUT"
The size alone does not prove that the metadata is usable. Run the relevant offline validation or inspection tool from your installed device-mapper tools, and review its result before presenting the destination to the target. The cache_repair manual names cache_check, cache_dump and cache_restore as related tools, but it does not define a universal validation sequence for every deployment.
After validation, keep the original metadata until the repaired destination has been tested. If the command returned 1, preserve its diagnostic output, do not activate the destination, and investigate the input, output capacity, permissions and offline state. Re-running against a newly chosen destination is safer than overwriting evidence from the failed attempt.
6. Return to service carefully
Only after the repaired metadata has passed your environment's checks should you use it with the device-mapper target. Follow the documented activation or replacement procedure for that host. This guide intentionally does not provide a generic activation command because the correct action depends on the device names, logical volumes and service manager in use.
There is no single undo option for a write to a metadata device. Recovery means retaining the original source, keeping a verified backup, and following the rollback procedure for the cache's owner. If you wrote to a regular output file, undo is straightforward: stop using it, preserve it for investigation, and select the original or another verified copy through your normal maintenance process. Do not delete the original merely because the repaired output exists.
Done means
- The installed version and command path were checked.
- The source was proven offline, not merely assumed to be idle.
- The input and output paths were reviewed and are different.
- A preallocated, sufficiently large destination was selected.
cache_repairreturned status0.- The output was inspected with the appropriate offline validation tool.
- The original metadata and a recovery path were retained.