Batch-convert Manual Pages to UTF-8 with man-recode
You will finish with a repeatable way to convert a set of manual-page files to one output encoding while keeping the originals available. The examples use the installed man-db 2.12.0, where man-recode can either create suffixed output files or replace each input after removing its compression extension.
The route
Jump straight to the step you need, or tick off Done means at the end.
Allow about ten minutes. You need the man-db package, readable manual-page files, and a writable destination directory. The commands normally run as your own user. Do not use sudo just because the source pages came from a system directory: copy them to a working directory first, or choose a writable output directory.
1. Check the installed command
Confirm the version and the option shape before putting the command in a script. This is a read-only check:
$ man-recode --version
man-recode 2.12.0
$ man-recode --help
Usage: man-recode [OPTION...]
-t CODE {--suffix SUFFIX | --in-place} FILENAME...
The important rule is that -t names the output encoding, and you must choose exactly one output mode: --suffix or --in-place. The installed package version can be checked independently with:
$ dpkg-query -W -f='${Package} ${Version}\n' man-db
man-db 2.12.0-4build2
Checkpoint: if man-recode is missing, stop here and use your normal package-management process. This guide does not install packages.
2. Inspect the inputs before converting
Make a working copy of the pages you intend to process. The input can be compressed, such as .gz, but the output created by man-recode is an uncompressed manual-page file:
$ mkdir -p "$HOME/man-recode-work"
$ cp /usr/share/man/man1/man-recode.1.gz "$HOME/man-recode-work/"
$ file "$HOME/man-recode-work/man-recode.1.gz"
/home/you/man-recode-work/man-recode.1.gz: gzip compressed data, ...
Replace the source path with your own file or a carefully selected list of files. Before a batch run, check both readability and the names that the shell will pass to the program:
$ find "$HOME/man-recode-work" -maxdepth 1 -type f -name '*.gz' -print
$ test -r "$HOME/man-recode-work/man-recode.1.gz" && echo readable
readable
Do not use an unreviewed wildcard against a broad directory. A wildcard can include generated output from an earlier run, or files that are not manual pages.
3. Create new files with a suffix
Use --suffix for the safest first conversion. It leaves the compressed input in place and appends your suffix after removing the input's compression extension:
$ man-recode --to-code=UTF-8 --suffix=.utf8 "$HOME/man-recode-work/man-recode.1.gz"
$ find "$HOME/man-recode-work" -maxdepth 1 -type f -printf '%f\n' | sort
man-recode.1.gz
man-recode.1.utf8
$ file "$HOME/man-recode-work/man-recode.1.utf8"
/home/you/man-recode-work/man-recode.1.utf8: troff or preprocessor input, ASCII text
The output name is man-recode.1.utf8, not man-recode.1.gz.utf8. The suffix is a naming choice, not a request to compress the result. Choose a suffix that cannot be confused with the original format, such as .utf8 or .converted.
For several known files, list them explicitly or use a narrow, reviewed glob:
$ man-recode -t UTF-8 --suffix=.utf8 \
"$HOME/man-recode-work/man-recode.1.gz" \
"$HOME/man-recode-work/other-page.1.gz"
Checkpoint: compare the file count and names before using the converted pages for installation. A successful exit status says that the conversion completed; it does not check whether your chosen pages belong in a particular package.
4. Understand how the input encoding is chosen
man-recode does not require you to provide a separate input encoding for each page. It first looks for an encoding declaration on the page's first line. A declaration can identify an encoding such as UTF-8 or ISO-8859-1; a first-line declaration that also names a preprocessor is supported. If there is no declaration, the program guesses from the file name.
This is a useful default for a mixed collection, but it is also the main source of surprising results. A misleading file name or a declaration that is not on the first line can lead to the wrong interpretation. Review the first line and the file naming convention when the output contains replacement characters or looks otherwise damaged:
$ zcat "$HOME/man-recode-work/man-recode.1.gz" | sed -n '1p'
."\ -*- coding: UTF-8 -*-
$ man-recode -d -t UTF-8 --suffix=.debug.utf8 \
"$HOME/man-recode-work/man-recode.1.gz"
guessed input encoding UTF-8 for /home/you/man-recode-work/man-recode.1.gz
The debug wording and path are environment-specific. Use --debug to see the input-encoding decision when diagnosing a page, not as a substitute for checking the resulting file.
5. Replace inputs only after a backup
--in-place writes the converted content back using the input base name, again removing a compression extension. This changes state and can destroy your only copy, so make a backup or use the suffix workflow instead:
$ cp --preserve=all \
"$HOME/man-recode-work/man-recode.1.gz" \
"$HOME/man-recode-work/man-recode.1.gz.bak"
$ man-recode -t UTF-8 --in-place \
"$HOME/man-recode-work/man-recode.1.gz"
$ find "$HOME/man-recode-work" -maxdepth 1 -type f -printf '%f\n' | sort
man-recode.1.gz.bak
man-recode.1
man-recode.1.utf8
After this operation, the converted file is man-recode.1; the original compressed name has been removed by the program. Check the result before deleting the backup:
$ file "$HOME/man-recode-work/man-recode.1"
$ man-recode -d -t UTF-8 --suffix=.verification \
"$HOME/man-recode-work/man-recode.1"
guessed input encoding UTF-8 for /home/you/man-recode-work/man-recode.1
Recovery is straightforward while the backup exists: remove the converted file and move or copy the backup back to its original name. Do not remove the backup as part of an unattended script until you have inspected the output and confirmed that any downstream installer accepts uncompressed pages.
6. Handle failures without hiding them
A missing or unreadable input produces a non-zero exit status. The default mode reports an error:
$ man-recode -t UTF-8 --suffix=.utf8 /path/to/missing-page.1
man-recode: can't open /path/to/missing-page.1
$ printf 'exit status: %s\n' "$?"
exit status: 1
--quiet suppresses conversion error messages, but it does not turn a failed conversion into a success. Keep it out of interactive troubleshooting and use it only when your surrounding script records and checks the exit status:
$ man-recode -q -t UTF-8 --suffix=.utf8 /path/to/missing-page.1
$ printf 'exit status: %s\n' "$?"
exit status: 1
If a page fails, check its path, permissions and first-line declaration. Do not respond by overwriting the source repeatedly or by running as root. A privilege change cannot repair an incorrect encoding declaration.
Done means
- You confirmed the installed man-db and
man-recodeversions. - You reviewed the exact input files and chose UTF-8 or another target encoding deliberately.
- You used
--suffixwhen preserving originals mattered. - You understood that compressed input produces an uncompressed output with its compression extension removed.
- You backed up inputs before using the destructive
--in-placemode. - You checked the output files and exit statuses instead of relying on a silent batch run.