Convert Chinese text safely with Perl's piconv
You will convert a Chinese text file between EUC-CN, GBK and UTF-8 using Perl's installed piconv, while keeping the source file intact. Allow about ten minutes if you know the source encoding. The examples use Perl 5.38.2 and the Encode module version 3.19 installed on this machine.
The route
Jump straight to the step you need, or tick off Done means at the end.
This is a byte-to-text conversion task, not a file rename. The source encoding must be correct. If you guess it, the command can complete successfully while producing mojibake.
1. Check the installed tools
Start with an unprivileged shell. You do not need sudo to read a text file or write a converted copy in a directory you own.
$ command -v perl piconv
/usr/bin/perl
/usr/bin/piconv
$ perl --version | sed -n '1,4p'
This is perl 5, version 38, subversion 2 (v5.38.2) built for x86_64-linux-gnu-thread-multi
$ perl -MEncode -e 'print Encode->VERSION, qq(\n)'
3.19
piconv --version is not a supported version query on this installation. It prints an unknown-option error followed by its usage text. Use perl --version and the module query above instead.
Checkpoint: you should have a path for piconv, a Perl version, and an Encode version. If the command is missing, stop and use your normal package-management process rather than copying a random script into /usr/local/bin.
2. Identify the source encoding
The local perlcn guide documents euc-cn as the name used for the traditional Unix Chinese encoding, and lists cp936 for code page 936, also called GBK. It also lists gb2312-raw, gb12345, iso-ir-165 and hz. These names describe different byte encodings, so do not select one because the filename happens to contain "gb".
Ask whoever produced the file, inspect the export settings, or use a trusted sample whose characters you can recognise. A UTF-8 file can be checked for a valid UTF-8 sequence, but validity alone does not prove that UTF-8 was intended. Preserve the original while testing.
$ input='/path/to/source.txt'
$ test -r "$input" && echo 'source is readable'
source is readable
$ file -- "$input"
/path/to/source.txt: Unicode text, UTF-8 text
The wording from file is useful evidence, not a substitute for knowing the producer's encoding. Replace the placeholder path before running the command. A path containing spaces is safe here because it is quoted.
3. Convert EUC-CN to UTF-8
Use -f for the source encoding and -t for the destination encoding. Redirect standard output to a new path. This is the practical form of the conversion shown by the installed manual:
$ piconv -f euc-cn -t utf8 < /path/to/file.euc-cn > /path/to/file.utf8
$ test -s /path/to/file.utf8 && echo 'converted file is non-empty'
converted file is non-empty
$ file -- /path/to/file.utf8
/path/to/file.utf8: Unicode text, UTF-8 text
The command reads the original through standard input and writes the converted bytes to standard output. It does not edit the input file. The utf8 spelling is the one used in the local guide; UTF-8 is commonly written with a hyphen in prose, but use the installed command's accepted encoding name in scripts.
Do not treat a non-empty output as proof that the conversion was right. Open it in an editor configured for UTF-8, or compare a few known Chinese characters with a trusted copy.
4. Convert UTF-8 to GBK
For a consumer or system that expects GBK, use the encoding name cp936. The local documentation says that GBK is also accepted for this code page, but cp936 makes the mapping explicit in a script:
$ piconv -f utf8 -t cp936 < /path/to/file.utf8 > /path/to/file.gbk
$ file -- /path/to/file.gbk
/path/to/file.gbk: Non-ISO extended-ASCII text
file may describe a GBK file differently on another release. The meaningful checks are that the command exits successfully, the destination is the expected size, and a consumer known to expect GBK reads the Chinese text correctly. Do not convert a file in place with a shell redirection such as piconv ... < file > file: the shell truncates the destination before piconv can read it.
5. Use a temporary output before replacing a file
When the destination name already exists, make the replacement conditional on a successful conversion. The temporary file should be in the same directory so that the final rename is on the same filesystem:
$ source='/path/to/file.euc-cn'
$ destination='/path/to/file.utf8'
$ temporary="${destination}.tmp"
$ piconv -f euc-cn -t utf8 < "$source" > "$temporary" \
&& test -s "$temporary" \
&& mv -- "$temporary" "$destination"
$ test -s "$destination" && echo 'replacement is ready'
replacement is ready
This changes state only at the final mv. If piconv fails, the && chain stops and the previous destination remains. Remove an incomplete temporary file after checking it is the one created by this command:
$ ls -l -- "$temporary"
$ rm -- "$temporary"
That rm is irreversible, so do not run it with a broad wildcard. If you replaced the wrong destination, recover from your backup, filesystem snapshot or version-controlled copy. The converter itself has no undo operation.
6. Handle conversion errors
A failure usually means the source name is wrong, the input contains bytes that the selected encoding cannot represent, or the file is not text in that encoding. First preserve the error and inspect the input without rewriting it:
$ piconv -f euc-cn -t utf8 < /path/to/file.euc-cn > /tmp/file-test.utf8
$ status=$?
$ printf 'piconv exit status: %s\n' "$status"
piconv exit status: 0
Use a temporary path under /tmp only for a disposable test. A non-zero status means you should not promote that output. Check the producer's encoding, then retry with the documented name that matches it. For an input declared as GBK, for example, try -f cp936 rather than -f euc-cn.
When the file contains characters absent from the destination encoding, conversion to GBK or EUC-CN may not be lossless. Keep UTF-8 as the archival or working copy, and make a separate legacy-encoding export for the system that requires it. Do not silently discard characters merely to make an export finish.
7. Know when a Perl script is the better option
piconv is convenient for whole-file conversion. A Perl program is clearer when only selected filehandles or strings should be decoded and encoded. The encoding module shown by the local guide can set the default encoding for standard input, output and error, but global defaults can surprise a larger program. For new code, keep encoding boundaries local and document them beside the filehandle.
$ perl -Mencoding=euc-cn,STDOUT,utf8 -pe1 < /path/to/file.euc-cn > /tmp/file.utf8
$ file -- /tmp/file.utf8
This compact command is useful for a one-off EUC-CN stream, but the explicit piconv -f ... -t ... form is easier to audit in a shell script because both directions are visible. The temporary output still needs checking before it replaces a useful file.
Done means
- You confirmed the installed Perl and Encode versions and found
piconv. - You identified the source encoding instead of guessing from the filename.
- You used
-fand-tin the correct direction. - You wrote to a new or temporary file, so the original remains available.
- You checked exit status, file presence and readable Chinese text before replacement.
- You kept UTF-8 when a legacy destination encoding could not represent every character.