Home / Alt manpages / perlko(1)

  • perlko(1)
  • User command
  • linux

Convert Korean Text Safely with Perl and piconv

You will convert a Korean text file between EUC-KR and UTF-8, while keeping the original available for recovery. The examples use the Perl 5.38.2 installation supplied by the perl-doc package on this machine. Allow about ten minutes if you know the source encoding; allow longer if you first need to identify it.

1. Check the installed tools

This guide assumes a readable input file and a shell account that can write in the working directory. No elevated privileges are needed for conversion. The relevant local manual is dated 14 September 2026 and identifies Perl as version 5.38.2. Check the paths before using a script or a scheduled job:

$ perl -v
This is perl 5, version 38, subversion 2 (v5.38.2)
$ command -v piconv
/usr/bin/piconv
$ piconv -l | grep -E '^(euc-kr|cp949|johab|utf8)$'
cp949
euc-kr
johab
utf8

Your piconv -l output may include aliases and additional encodings. The useful check is that the names you plan to use are available. The manual lists euc-kr, cp949, johab, iso-2022-kr and ksc5601-raw as Korean encodings supported by the Encode module.

2. Preserve the source file

Do not guess the input encoding. A file called old.txt gives no reliable indication whether its bytes are EUC-KR, CP949 or UTF-8. Confirm the format from the producing system, export settings or archive notes. A wrong source encoding can produce replacement characters or unreadable text even when the command exits successfully.

Make a copy before the first conversion. Use a separate destination name so shell redirection cannot truncate the source:

$ cp --preserve=all /path/to/file.euc-kr /path/to/file.euc-kr.bak
$ test -r /path/to/file.euc-kr && echo readable
readable

The backup is your recovery point. Keep it until the converted file has been opened and checked in the application that will consume it. Removing the backup is irreversible, so do that later and deliberately.

3. Convert EUC-KR to UTF-8

The local manual gives two equivalent styles. The Perl command applies the euc-kr encoding to input and writes UTF-8 to standard output. Redirect that output to a new file:

$ perl -Mencoding=euc-kr,STDOUT,utf8 -pe1 < /path/to/file.euc-kr > /path/to/file.utf8
$ file /path/to/file.utf8
/path/to/file.utf8: Unicode text, UTF-8 text

The exact file description depends on its version, so treat it as a hint rather than proof. Check the result in a UTF-8-aware editor or application, and compare the number of records with the source if the format has one record per line:

$ wc -l /path/to/file.euc-kr /path/to/file.utf8
  120 /path/to/file.euc-kr
  120 /path/to/file.utf8
  240 total

Do not rely on byte counts to match. Korean characters can occupy different numbers of bytes in different encodings.

4. Use piconv for the same conversion

piconv is the bundled Perl utility for this job. Its -f option names the source encoding and -t names the destination encoding:

$ piconv -f euc-kr -t utf8 < /path/to/file.euc-kr > /path/to/file.utf8
$ test -s /path/to/file.utf8 && echo conversion-produced-a-non-empty-file
conversion-produced-a-non-empty-file

Use the piconv form when it makes a conversion script easier for another operator to read. The direction matters. To create EUC-KR for an older system from UTF-8 input, reverse the two names and use a different destination:

$ piconv -f utf8 -t euc-kr < /path/to/file.utf8 > /path/to/file.euc-kr.new
$ file /path/to/file.euc-kr.new

Some Unicode characters cannot be represented by the target encoding. Treat warnings or conversion errors as a reason to inspect the data and destination requirements, not as permission to discard the original.

5. Make replacement safe

Never put > /path/to/file.utf8 in a blind batch if that file may already be useful. Redirection truncates the destination before the converter has finished. Write a temporary output, inspect it, then replace the old file only when you intend to:

$ piconv -f euc-kr -t utf8 < /path/to/file.euc-kr > /path/to/file.utf8.new
$ test -s /path/to/file.utf8.new
$ mv /path/to/file.utf8.new /path/to/file.utf8
$ ls -l /path/to/file.utf8

If conversion fails, leave the existing destination alone and remove the incomplete .new file after checking the error. If the move has already happened and the result is wrong, restore the destination from file.euc-kr.bak by rerunning the conversion, or copy a separately preserved known-good output over it. Keep the source backup until that check is complete.

6. Keep source code and data encodings separate

The manual distinguishes Perl source text from input and output data. Store source as UTF-8 and use use utf8; when the program itself contains non-ASCII characters. That pragma does not detect the encoding of an external file. For file handles, choose an explicit encoding and document the choice. Mixing a UTF-8 source file, EUC-KR input and an unspecified terminal is a common way to misdiagnose a display problem as data corruption.

Also watch for double encoding: converting bytes that have already been decoded, or decoding the same bytes twice, changes the data. Test with representative Korean text and preserve the original bytes while investigating.

Done means

  • The installed Perl and piconv versions and paths were checked.
  • The source encoding was confirmed rather than inferred from the filename.
  • The original file remains available as a backup.
  • The output was written to a new file, checked for content and opened in its target application.
  • The conversion direction and target encoding are recorded for the next operator.