Home / Alt manpages / perltw(1)

  • perltw(1)
  • User command
  • linux

Convert Big5 Text Safely with Perl's piconv

You will convert a Taiwan Chinese text file from Big5 to UTF-8, check which encoding names this Perl installation accepts, and keep the original file untouched. The examples use piconv, Perl's bundled character-conversion utility, rather than guessing from a filename. Allow about fifteen minutes, plus time to inspect the converted text.

This guide is for a Linux system with the perl-doc package and Perl 5.38.2, as installed here. You need a readable input file and a writable working directory. The commands are ordinary user commands. Nothing here needs sudo.

1. Check the installed tools

Confirm that the command resolves to the binary you expect and record the Perl package version. This is a read-only checkpoint:

$ command -v perl piconv
/usr/bin/perl
/usr/bin/piconv
$ dpkg-query -W -f='${Package} ${Version}\n' perl-doc perl-base
perl-doc 5.38.2-3.2ubuntu0.6
perl-base 5.38.2-3.2ubuntu0.6

piconv is supplied with Perl. Its interface reads standard input or named files and writes converted data to standard output. That output design is useful for a new destination, but shell redirection can still truncate an existing file, so choose output names deliberately.

2. Identify the encoding names

The name Big5 is not the only spelling you may encounter. The local Encode::TW documentation describes big5-eten, big5-hkscs and cp950, among others. Ask this installation to resolve aliases instead of relying on memory:

$ piconv -r big5
big5-eten
$ piconv -r cp950
cp950

Use the name that matches the data's origin. big5-eten is the normal Big5 encoding with ETen extensions. cp950 is Microsoft's code page variant. They are related, but they are not interchangeable for every byte sequence. If you do not know which one produced the file, make a copy and test the likely choices against known text rather than overwriting anything.

Checkpoint: list the encodings available to the installed Perl if you need to investigate an unfamiliar label:

$ piconv -l | grep -iE 'big5|cp950|utf-?8'
big5-eten
big5-hkscs
cp950
utf8

The list can vary with Perl and its installed modules. An alias that is not listed may still resolve, so piconv -r NAME is the more direct test.

3. Convert Big5 to a new UTF-8 file

Replace INPUT.big5 with the path to your file. This command reads the source and creates OUTPUT.txt:

$ piconv -f big5 -t utf8 INPUT.big5 > OUTPUT.txt

-f selects the input encoding and -t selects the output encoding. The short options also have long forms, --from and --to. Both are worth spelling out in a script when the encoding choice needs to be obvious.

Check that the destination is non-empty and that it is valid UTF-8:

$ test -s OUTPUT.txt && echo 'output is non-empty'
output is non-empty
$ file OUTPUT.txt
OUTPUT.txt: Unicode text, UTF-8 text

Your file wording may differ. The useful result is that it identifies UTF-8 or Unicode text, and that the size is plausible for the input. Open the result in a UTF-8-aware editor or print a small section to inspect the characters.

4. Convert the other direction

To create Big5 from a UTF-8 source, reverse the two encoding arguments and write a separate destination:

$ piconv -f utf8 -t big5 UTF8_INPUT.txt > BIG5_OUTPUT.txt
$ file BIG5_OUTPUT.txt
BIG5_OUTPUT.txt: Non-ISO extended-ASCII text

Not every Unicode character has a Big5 representation. When a character cannot be represented, the default result can contain a replacement character or another fallback, depending on the conversion path. Use a small known sample first, and inspect the result before processing an archive. For a diagnostic conversion that checks the stream, add -c:

$ piconv -c -f utf8 -t big5 UTF8_INPUT.txt > BIG5_OUTPUT.txt
$ printf 'conversion status: %s\n' "$?"
conversion status: 0

A non-zero status is a reason to stop and investigate the input and target encoding. Do not treat a file produced before an error as complete.

5. Avoid shell redirection traps

Never use an existing source or valuable destination as the output path. The shell opens the destination before piconv runs, so a command such as piconv ... > INPUT.big5 can erase the input immediately.

For a replacement workflow, write a temporary file in the same directory, verify it, then move it into place. The final mv changes the directory entry and is the deliberate, irreversible step:

$ piconv -f big5 -t utf8 INPUT.big5 > INPUT.big5.new
$ test -s INPUT.big5.new
$ file INPUT.big5.new
INPUT.big5.new: Unicode text, UTF-8 text
$ mv -- INPUT.big5.new INPUT.txt

If conversion fails, remove only the incomplete INPUT.big5.new after checking the error; the source remains available. If you moved the wrong file, stop rather than guessing. Restore it from your backup or rename it back with mv -- INPUT.txt INPUT.big5 only when you have confirmed those exact paths.

6. Use Perl directly when the conversion is part of a script

The underlying API is the Encode module. It separates decoding bytes into Perl's character strings from encoding those strings back into bytes. This example follows the documented Big5-to-UTF-8 pipeline:

$ perl -MEncode -pe '$_ = encode("utf8", decode("big5", $_))' < INPUT.big5 > OUTPUT.txt

Use piconv for a one-off or simple pipeline. Use the module when a program must choose an encoding, handle errors, or process several files. Keep the source and destination encodings explicit. A filename, locale or terminal setting does not prove how the bytes inside an old file are encoded.

Done means

  • perl and piconv resolve to the intended installation.
  • The input encoding was chosen from the file's provenance or tested against known text.
  • The conversion used explicit -f and -t values.
  • The output is non-empty and identified as the expected text encoding.
  • The original file was not overwritten until a verified replacement was ready.