uconv converts text between encodings and tells you exactly which characters the destination cannot represent. You will also finish with a way to check an input file without creating output at all. The examples use uconv 74.2 from the Debian package icu-devtools version 74.2-1ubuntu3.1.
Allow about ten minutes. You need a shell, a readable text file, and the ICU development tools package. None of the commands below needs root. Work on a copy when the input is valuable: conversion can lose information, and uconv does not edit an input file in place unless you deliberately redirect output back over it.
Start by confirming which executable and package version you are using. This is a read-only checkpoint:
$ command -v uconv
/usr/bin/uconv
$ uconv --version
uconv v2.1 ICU 74.2
$ dpkg-query -W -f='${Package} ${Version}\n' icu-devtools
icu-devtools 74.2-1ubuntu3.1
The manpage identifies the utility as ICU 74.2. Encoding names belong to ICU, so do not assume that every name accepted by GNU iconv will work here.
List all encodings when you need to choose one:
$ uconv --list | head
The list is long and its order is not a useful contract. Check one candidate directly before putting it in a script:
$ uconv --list-code utf-8
UTF-8
For the machine default, ask explicitly:
$ uconv --default-code
UTF-8
A missing or invalid encoding produces an error instead of a conversion. Keep the exact spelling from --list or --list-code.
Use -f for the input encoding, -t for the destination, and -o for a new file. This example leaves input.txt untouched:
$ uconv -f utf-8 -t iso-8859-1 -o converted.txt input.txt
$ file converted.txt
converted.txt: ISO-8859 text
The conversion passes through Unicode. That makes the two encodings explicit, but it does not make an unrepresentable character safe. With the default callback, a character that cannot be converted to the destination stops the command with an error.
Checkpoint: compare the files before replacing anything:
$ wc -c input.txt converted.txt
$ cmp -s input.txt converted.txt; printf 'cmp status: %s\n' "$?"
cmp status: 1
A non-zero cmp status is expected when the byte encoding changed. It is not by itself proof that the text is wrong. Inspect representative content and keep the original until the result is accepted.
To check whether a file can be read as a particular encoding, send converted output to /dev/null. The -c option skips invalid characters, so this is a validation command rather than a faithful conversion:
$ uconv -f utf-8 -c input.txt >/dev/null
$ printf 'exit status: %s\n' "$?"
exit status: 0
Status 0 means the command completed. It does not mean that every byte was preserved, because -c deliberately omits invalid data. For a strict check, omit -c and use the default stop callback:
$ uconv -f utf-8 input.txt >/dev/null
$ printf 'exit status: %s\n' "$?"
exit status: 0
With damaged input, the strict command exits non-zero. The reported offset is the first invalid byte that ICU encounters. For multi-byte encodings, that can differ from the position reported by GNU iconv.
Choose the policy deliberately. --callback stop is the default and preserves the evidence of a conversion problem. -c or --to-callback skip loses the offending output. --to-callback escape-xml-dec keeps a visible numeric reference, which is useful when the result will be consumed as HTML:
$ printf 'café €\n' | uconv -f utf-8 -t us-ascii --callback escape-xml-dec
café €
The numeric references are output text, not automatic HTML markup. Encode the result for the context in which it will be used. Other useful callbacks include substitute, escape-c, escape-java, escape-unicode, and stop. List the installed command's accepted names with uconv --help.
Safety boundary: Do not use -c in a migration merely to obtain a successful exit status. First decide whether silent data loss is acceptable, then record that decision beside the conversion job.
The -x option runs an ICU transliterator between the input and output conversion. A small, inspectable example turns a character into its Unicode name:
$ printf 'カ' | uconv -f utf-8 -x 'hex-any; any-name'
\N{KATAKANA LETTER KA}
Transliterator identifiers are case-sensitive enough to deserve a check. List the available identifiers before using one in a script:
$ uconv --list-transliterators | grep -E '^(Any-Name|Hex-Any|NFKC)$'
Identifiers can be chained with semicolons. ICU's documentation also describes normalisation and rule-based transforms, but test the exact spelling on the installed version. In this environment, the manpage's illustrative ::nfkc form is rejected by the installed command; uconv --list-transliterators is the authority for what this binary accepts.
A Unicode signature is a byte-order mark, commonly called a BOM. Add or remove it only when the receiving format requires that change:
$ uconv -f utf-8 -t utf-8 --remove-signature -o without-bom.txt input.txt
$ uconv -f utf-8 -t utf-8 --add-signature -o with-bom.txt input.txt
$ file without-bom.txt with-bom.txt
These commands create two new files. If you later need to replace a production file, make a backup first and use a temporary output in the same directory before an explicit rename. A rename is a state change and can affect readers immediately; do not run it as an unreviewed batch operation. Recovery is to restore the backup with the same ownership and permissions.
-f uses the default encoding, which is not a safe guess for an unknown file. Identify the source encoding first.-t also uses the default encoding. Set both sides in repeatable jobs.-o with a path that already exists can replace that file. Choose a fresh output name while testing.--fallback permits compatibility mappings that are not strict character matches. Leave the default --no-fallback in place unless the receiving format calls for fallback mappings.uconv version.