Home / Alt manpages / perlunitut(1)

  • perlunitut(1)
  • User command
  • linux

Perl Unicode Text: Decode Input, Encode Output with Encode

You will finish with a small Perl program that receives UTF-8 bytes, turns them into characters, processes those characters, and emits UTF-8 bytes again. The boundary is explicit: decode on the way in and encode on the way out. That keeps character counts and byte counts from being mixed up.

Allow about 15 minutes. You need Perl and the core Encode module. This guide follows the installed perlunitut(1) tutorial from perl-doc 5.38.2. It assumes you already understand that Unicode is a character set, UTF-8 is an encoding, and bytes and characters are different things.

1. Check the Perl and Encode versions

Start with ordinary, read-only checks. No elevated privileges are needed:

$ perl -v
$ perl -MEncode -e 'print Encode->VERSION, "\n"'

The first command reports the installed interpreter. On the machine used for this guide it reports Perl 5.38.2. The second confirms that Encode can be loaded. The exact module version can differ from the interpreter version, so keep both checks when diagnosing a different host.

Checkpoint: if the second command fails with a missing module error, stop there. Do not work around it by treating raw bytes as characters. Install or repair the normal Perl package through your system's package-management process, then rerun the check.

2. Keep the three stages visible

The useful model is short enough to keep beside your editor:

  1. Receive binary data and decode it with the encoding you actually received.
  2. Process the resulting text as characters.
  3. Encode the finished text to the encoding required by the destination.

UTF-8 is a sensible standard when you control both ends, but it is not a synonym for Unicode. If a file or protocol says ISO-8859-1 or Windows-1251, pass that actual name to decode. Decoding with the wrong label can produce plausible-looking but incorrect text.

Binary data such as a PNG is not text. Keep it as bytes and do not send it through a text decoder. The workflow below is for text data only.

3. Decode input before counting or editing it

Add the two functions from Encode, then decode the bytes you received:

use strict;
use warnings;
use Encode qw(encode decode);

my $input_bytes = "caf\xC3\xA9\n";
my $text = decode('UTF-8', $input_bytes);

printf "characters: %d\n", length $text;
print "text: $text";

The byte sequence in $input_bytes is UTF-8 for cafe with an acute e, followed by a newline. After decoding, length counts characters in the text string. Run it as a normal user:

$ perl unicode-step.pl
characters: 5
text: café

The displayed accent may render differently in your terminal, but the character count should be 5: four letters plus the newline. The important operation is the conversion, not the terminal's font.

Do not call length on the original bytes when you mean character count. The UTF-8 encoding of some characters occupies multiple bytes, so a byte count and a character count are not interchangeable.

4. Process text as characters

Once the value has been decoded, normal string operations work on text. This example removes the newline, replaces a word and reports a character count:

chomp $text;
$text =~ s/cafe/café/;
my $character_count = length $text;
printf "processed: %s\ncharacters: %d\n", $text, $character_count;

For real input, choose operations that match the data's intended characters. A visible glyph can be made from more than one Unicode code point, and a user-perceived grapheme is not always the same as either a code point or a byte. This tutorial's practical boundary still matters: do not use byte operations to answer a character question.

Checkpoint: keep the value as text while substitutions, regular expressions, and character counts are happening. There should be no call to encode in the middle of this stage.

5. Encode once, immediately before output

When the destination expects UTF-8, convert the finished text to bytes at the output boundary:

my $output_bytes = encode('UTF-8', "$text\n");
my $byte_count = length $output_bytes;

print "Content-Type: text/plain; charset=UTF-8\n";
print "Content-Length: $byte_count\n\n";
print $output_bytes;

After encode, $output_bytes is a byte string. Its length is therefore the number of bytes to send, which is the quantity a byte-oriented protocol means by Content-Length. The content type tells the receiver which character encoding to use when decoding those bytes.

For a file, the same rule applies. Open or obtain the output destination according to the surrounding program, encode the completed text, then write the encoded bytes. Do not calculate a byte length before encoding and assume it remains valid.

6. Make the encoding contract fail loudly

decode must be given the source encoding. If the input is not valid for that encoding, investigate the producer or its documented protocol rather than guessing a different one. A guessed fallback can silently corrupt data.

The same caution applies to legacy encodings. An encoding such as ISO-8859-1 cannot represent every Unicode character. Converting text to a narrower destination can lose characters. Check the destination's supported encodings before choosing one, and retain the original text until the conversion has been verified.

Do not silence warnings or add an arbitrary encoding merely to make an error disappear. If a byte stream is actually binary, keep it binary. If it is text, establish its encoding at the input boundary.

7. Avoid the common boundary mistakes

  • Do not confuse Unicode with UTF-8. Unicode supplies characters; UTF-8 supplies bytes for storing or transmitting them.
  • Do not decode the same bytes twice. Decode once when they enter the text-processing part of the program.
  • Do not encode a text string, edit it as though it were characters, and then encode it again.
  • Do not decode images, archives or other binary payloads as text.
  • Do not guess an input encoding from a successful command. Obtain it from the file format or protocol.
  • Do not use a character count where a protocol requires a byte count. Encode first, then measure.

There is no state to undo in these examples: they only transform values in a short-lived Perl process. If you adapt the pattern to overwrite a file or send a production response, test with a new output path or a safe staging destination first. A shell redirection to an existing file can truncate it before Perl reports an error.

Done means

  • Encode loads on the installed Perl 5.38.2 interpreter.
  • Incoming text is decoded with the source encoding before character operations.
  • Processing uses text strings, so character counts are taken before output conversion.
  • Finished text is encoded once at the destination boundary.
  • Any byte count is measured after encoding, and the declared charset matches the bytes sent.
  • Binary payloads and unknown encodings are kept out of the text path until their format is established.