Compress and restore files safely with pigz
You will compress a file with the installed pigz 2.8 package, verify the resulting gzip stream, and restore the original without losing your first copy. The same command can use multiple processors for compression, while decompression remains fundamentally single-threaded. Allow about ten minutes for the first run. You need a shell, a readable input file and write access to its directory or to a separate working directory.
The route
Jump straight to the step you need, or tick off Done means at the end.
This guide changes files in the examples. The default compression mode removes the input after a successful operation, so start with -k while learning. Do not use sudo unless the files themselves require elevated access; changing ownership or permissions to make a personal archive writable is a separate decision.
1. Check the installed command
Confirm which executable your shell will run and record its version:
$ command -v pigz
/usr/bin/pigz
$ pigz --version
pigz 2.8
The local package is pigz 2.8-1. The manual page is dated 19 August 2023 and documents gzip, zlib and single-entry zip output. The examples here use the default gzip format, which produces a .gz file.
Checkpoint
If command -v pigz prints nothing, stop and install the package through your normal system-management process. Do not substitute an unverified binary into a backup workflow.
2. Compress a file while keeping the original
Make a small test file in a directory you control, or replace the path below with an existing file. The -k option keeps the input:
$ printf '%s\n' 'pigz test data' > sample.txt
$ pigz -k sample.txt
$ ls -l sample.txt sample.txt.gz
-rw-r--r-- 1 user user 15 ... sample.txt
-rw-r--r-- 1 user user 35 ... sample.txt.gz
The exact size, owner and timestamp vary. The useful result is that both names exist and the command returned status 0. Compression uses the default level -6. By default pigz uses the number of online processors; add -p 1 when you specifically need compression without worker threads.
Without -k, a successful file compression deletes the original input. That is a destructive change. If you did run it accidentally, do not create a new file with the same name and assume it is the original; restore from a known backup or decompress the archive after checking it.
3. Verify the archive before relying on it
Ask pigz to test the compressed input without writing a restored file:
$ pigz -t sample.txt.gz
$ printf 'status: %s\n' "$?"
status: 0
No output is normal for a successful test. A non-zero status means the stream could not be verified, so treat the archive as damaged or unsuitable until you investigate. Keep the original sample.txt until this check and a content check have passed.
For a second check, compare the uncompressed bytes with the original. This reads the archive and writes nothing:
$ pigz -dc sample.txt.gz | cmp -s - sample.txt
$ printf 'comparison: %s\n' "$?"
comparison: 0
The -d option decompresses, -c sends the result to standard output, and cmp compares that stream with the original. A zero comparison status means the bytes match. The pipeline is safe for this check because it does not replace either file.
4. Restore the original file
When the archive is in the same directory, the short form unpigz is equivalent to using pigz -d. Use -k if you want to retain the compressed archive:
$ rm sample.txt
$ unpigz -k sample.txt.gz
$ ls -l sample.txt sample.txt.gz
-rw-r--r-- 1 user user 15 ... sample.txt
-rw-r--r-- 1 user user 35 ... sample.txt.gz
The rm above is intentionally shown only to model a directory containing the archive and no restored file. It is irreversible, so do not copy it blindly if your original still matters. If sample.txt already exists, unpigz normally refuses to overwrite it. Use a different output path with standard output instead:
$ unpigz -c sample.txt.gz > restored-sample.txt
$ cmp -s restored-sample.txt sample.txt
$ printf 'restored bytes match: %s\n' "$?"
restored bytes match: 0
Shell redirection truncates an existing destination before unpigz starts. Choose a new destination, or make a backup first. If the command fails part-way through, remove only the incomplete destination after checking its path; the gzip archive remains unchanged.
5. Stream data without creating an intermediate file
With no input name, or with -, pigz reads standard input. With -c, it writes compressed data to standard output. This makes it suitable for a pipeline:
$ tar -cf - /path/to/directory | pigz -c > directory.tar.gz
$ pigz -t directory.tar.gz
$ pigz -dc directory.tar.gz | tar -tf - | sed -n '1,5p'
Use a destination name that does not already contain valuable data. The pipeline does not give pigz a normal filename to store in the gzip header, and an interrupted pipeline can leave a truncated archive. Check the exit status of each stage when scripting, rather than treating the existence of directory.tar.gz as proof of success.
For a stream that must be sent elsewhere, keep the archive on standard output and redirect it only at the final boundary. Avoid mixing diagnostic text into a binary stream: do not use verbose output on the same standard output that another program will consume.
6. Choose speed, size and recovery behaviour
The compression level ranges from -1, fastest and usually larger, through the default -6 to -9, slower and usually smaller. Level -11 uses the zopfli algorithm and can take much longer for only a modest size improvement. Select a level for the job rather than assuming the slowest setting is automatically best:
$ pigz -k -1 sample.txt
$ pigz -k -9 sample.txt
$ ls -l sample.txt.gz
Do not run those two commands unchanged against the same destination in a real archive directory: the second command may meet an existing .gz file. Use separate output directories or names when comparing settings, and keep the command's output files distinct.
Normal pigz blocks use the preceding 32 KiB as a preset dictionary, which helps compression. Add -i when independently decompressible blocks are more valuable than that compression benefit. The option is for damage recovery or random-access designs, not a general speed switch. The input block size is 128 KiB by default and can be changed with -b; the number of compression processes can be limited with -p.
7. Inspect files and handle awkward names
Use -l to list information about a compressed input and -v for more verbose messages:
$ pigz -lv sample.txt.gz
$ pigz -l sample.txt.gz
Exact columns and ratios depend on the installed build and input. Use the output as a quick inspection, then use -t for the integrity decision. A filename beginning with - can be treated as an option unless you place -- before it:
$ pigz -k -- -odd-name.txt
Environment variables can also alter behaviour: the manual says options are read from GZIP and then PIGZ before command-line options. When a script behaves unexpectedly, inspect those variables before blaming the file or adding -f. Avoid -f unless overwriting an existing archive or processing a link is an explicit, reviewed decision.
Done means
pigz --versionidentified the executable you intended to use.- The archive has the expected
.gzname andpigz -treturns status 0. - A byte comparison or an equivalent application check confirms the restored data.
- The original or a separate backup still exists when recovery matters.
- You know whether the next operation will delete, overwrite or preserve a file.