Home / Alt manpages / perljp(1)

  • perljp(1)
  • User command
  • linux

Convert Japanese Text with Perl Encode and piconv

You will finish with a repeatable way to convert Japanese text between UTF-8, EUC-JP and Shift_JIS, using the Perl 5.38.2 tools installed with Ubuntu's perl-doc package. The key rule is simple: decode incoming bytes into Perl characters, then encode characters for the destination.

Allow about fifteen minutes. You need a shell, Perl, and the piconv utility. Check the package first; all examples are ordinary user commands and should not need sudo. Work on copies of valuable files.

1. Check the installed tools

The local perljp(1) guide documents Unicode support introduced in Perl 5.8 and lists Japanese encodings including euc-jp, shiftjis and utf8. Confirm the actual interpreter and converter before relying on an example:

$ perl -v | sed -n '1,8p'
This is perl 5, version 38, subversion 2 (v5.38.2)

$ command -v piconv
/usr/bin/piconv

$ piconv -r shift_jis
shiftjis

Checkpoint: the version and paths can differ on another host. Use the encoding names accepted by that installed Encode module, rather than guessing from a filename extension.

2. Make a UTF-8 test file

Start with a small known-good UTF-8 file. This command writes Japanese text followed by a newline:

$ printf '\343\201\202\343\201\204\n' > japanese-utf8.txt
$ file japanese-utf8.txt
japanese-utf8.txt: Unicode text, UTF-8 text

The byte escapes in the example are deliberate. They let the shell create the same UTF-8 bytes without depending on the terminal locale. In normal work, a UTF-8 editor can create the source file instead.

Do not confuse a successful conversion with a correct conversion. Keep the original file until the destination has been checked. Shell redirection with > truncates an existing destination before the program runs, so choose a new name while testing.

3. Convert a file with piconv

piconv is the practical command-line front end to Perl's Encode support. Give it the input encoding with -f and the output encoding with -t:

$ piconv -f utf8 -t euc-jp japanese-utf8.txt > japanese-euc-jp.txt
$ piconv -f utf8 -t shiftjis japanese-utf8.txt > japanese-shiftjis.txt
$ file japanese-euc-jp.txt japanese-shiftjis.txt
japanese-euc-jp.txt:      Non-ISO extended-ASCII text
japanese-shiftjis.txt:   Non-ISO extended-ASCII text

The exact file descriptions vary by version and locale. A stronger round-trip check is to convert each result back to UTF-8 and compare the bytes with the original:

$ piconv -f euc-jp -t utf8 japanese-euc-jp.txt > roundtrip-euc.txt
$ piconv -f shiftjis -t utf8 japanese-shiftjis.txt > roundtrip-shiftjis.txt
$ cmp --silent japanese-utf8.txt roundtrip-euc.txt && echo 'EUC-JP round trip: OK'
EUC-JP round trip: OK
$ cmp --silent japanese-utf8.txt roundtrip-shiftjis.txt && echo 'Shift_JIS round trip: OK'
Shift_JIS round trip: OK

If cmp reports a difference, stop and inspect the source encoding, destination encoding and error output. Do not delete the source to make the mismatch disappear.

4. Convert a string without creating a file

For a one-off value, use -s. This is useful for checking an encoding name or inspecting the result:

$ piconv -f utf8 -t euc-jp -s '日本語' | od -An -tx1
 c6 fc cb dc b8 ec

The displayed hexadecimal bytes are the EUC-JP representation of the three characters. They are not meant to be read as Japanese in the terminal. Use piconv -l to list encodings supported by the installed module, or piconv -r NAME to resolve an alias to its canonical name.

5. Use Encode inside a Perl program

Use Encode::decode at the boundary where bytes enter the program and Encode::encode where bytes leave it. This example reads UTF-8 from standard input and writes EUC-JP:

$ printf '\343\201\202\343\201\204\n' |
  perl -MEncode=decode,encode -ne \
  'print encode("euc-jp", decode("UTF-8", $_))' | od -An -tx1
 a4 a2 a4 a4 0a

$_ contains the input line's bytes here. Decode turns those bytes into Perl's internal character string; encode turns that string into the requested output bytes. Keeping those operations explicit prevents a locale or filehandle default from silently deciding the conversion.

For a larger script, put the conversion in a small function and name the boundary encodings. That makes a later change from Shift_JIS to UTF-8 visible in code review.

6. Account for the old encoding example

The installed perljp(1) page also shows a -Mencoding=... approach for migrating older JPerl scripts. That section describes historical Perl behaviour. On this Perl 5.38.2 installation, the pragma is rejected with The encoding pragma is no longer supported.

Use explicit Encode calls or piconv for new work. If an old script depends on the pragma, preserve a copy, run its tests, and plan a small migration rather than changing its shebang blindly. The failure is a compatibility signal, not a reason to install a second Perl over the system interpreter.

7. Handle bad input carefully

A conversion can fail because the input is not valid in the encoding you named, or because a character cannot be represented in the destination encoding. First rerun the command with a different, verified source encoding. Do not solve uncertainty by trying several encodings and keeping whichever output looks plausible.

For production conversion, write to a temporary destination in the same directory, check it, then replace the old file in a separate deliberate step. For example:

$ piconv -f utf8 -t euc-jp japanese-utf8.txt > japanese-euc-jp.txt.new
$ piconv -f euc-jp -t utf8 japanese-euc-jp.txt.new > /tmp/japanese-check.txt
$ cmp --silent japanese-utf8.txt /tmp/japanese-check.txt && echo 'checked'
checked
$ mv japanese-euc-jp.txt.new japanese-euc-jp.txt

mv in this example replaces the destination. Do not run it until the round trip is correct. If the check fails, remove only the .new file and retain the previous destination. The original UTF-8 source remains the recovery copy.

Done means

  • You confirmed the local Perl version and piconv path.
  • You named both the source and destination encodings explicitly.
  • A file conversion passed a byte-for-byte UTF-8 round-trip check.
  • You can use Encode::decode and Encode::encode at program boundaries.
  • You know the old encoding pragma example in perljp(1) does not run on Perl 5.38.2.
  • You keep the source and use a checked temporary output before replacing a useful file.