Make Perl Unicode Text Predictable with Explicit Encoding
You will finish with a small Perl program that reads UTF-8 text as characters, writes UTF-8 text deliberately, and reports the difference between character count and byte count. The examples target Perl 5.38.2, the version installed on this machine. Allow about fifteen minutes if Perl is already installed. You need a shell and a text editor; no elevated privileges are required.
The route
Jump straight to the step you need, or tick off Done means at the end.
Checkpoint: the central rule is simple. Decode bytes when they enter the program, work with characters, then encode characters when they leave it. A Unicode-aware string is not the same thing as UTF-8 bytes on disk or on a pipe.
1. Check the installed Perl version
Start with a read-only version check. This tells you which local behaviour the examples describe:
$ perl -v
This is perl 5, version 38, subversion 2 (v5.38.2) built for x86_64-linux-gnu-thread-multi
$ perl -MConfig -e 'print "$Config{version}\n"'
5.38.2
The perlunicode manual installed with this Perl is dated 14 September 2026. Later Perl releases may refine Unicode details, so use the local manual when behaviour differs from this guide.
2. Declare Unicode rules and UTF-8 source
Put use v5.12 or a later version declaration near the top of new code. The manual says this automatically enables the unicode_strings feature. That feature changes how strings are interpreted; it does not configure filehandles.
If the source file itself contains non-ASCII characters, add use utf8. This tells Perl to interpret string literals, regular expressions and identifiers in the source file as UTF-8. It is not a replacement for an I/O encoding layer.
#!/usr/bin/perl
use v5.12;
use utf8;
use strict;
use warnings;
my $word = 'café';
print "$word\n";
Save the file as UTF-8 and run it normally:
$ perl unicode-demo.pl
café
Checkpoint: if the output looks wrong, inspect the terminal locale and the file encoding before changing the string. use utf8 only governs how Perl reads the program source.
3. Make file I/O explicit
Use an :encoding(UTF-8) layer for text files. This decodes UTF-8 input into Perl characters and encodes output characters as UTF-8 bytes. The example creates a test file in the current directory and then reads it back.
#!/usr/bin/perl
use v5.12;
use utf8;
use strict;
use warnings;
open my $in, '<:encoding(UTF-8)', 'message.txt' or die "read message.txt: $!";
open my $out, '>:encoding(UTF-8)', 'copy.txt' or die "write copy.txt: $!";
while (my $line = <$in>) {
print {$out} $line or die "write copy.txt: $!";
}
close $in or die "close message.txt: $!";
close $out or die "close copy.txt: $!";
Create a known UTF-8 input without relying on the terminal:
$ printf 'café\n東京\n' > message.txt
$ perl copy-utf8.pl
$ cmp -- message.txt copy.txt
$ echo $?
0
Both printf and Perl are ordinary user commands. The redirection operator overwrites copy.txt, so choose a disposable path or remove the destination before testing. Do not use sudo to work around a directory permission problem unless the directory is intentionally administrator-only.
4. Measure characters and bytes separately
Perl string functions such as length normally count characters, not the number of UTF-8 bytes used to store them. Use the bytes pragma only for a byte-oriented measurement, or encode a copy explicitly with Encode.
#!/usr/bin/perl
use v5.12;
use utf8;
use strict;
use warnings;
use Encode qw(encode);
my $text = 'café';
my $utf8 = encode('UTF-8', $text);
print "characters: ", length($text), "\n";
print "UTF-8 bytes: ", length($utf8), "\n";
Run it and check the expected relationship:
$ perl count-text.pl
characters: 4
UTF-8 bytes: 5
The accented character is one character but takes two UTF-8 bytes. The exact byte count changes with the text. If a protocol specifies a byte limit, measure the encoded value before sending it. If a user-facing limit is in characters, do not substitute a byte count.
5. Decode data from a byte-oriented interface
Some APIs and older extensions return byte strings without documenting Unicode support. Treat that boundary as untrusted. Decode known UTF-8 bytes with Encode::decode, and encode the result again when passing it to an interface that expects raw UTF-8.
use Encode qw(decode encode);
my $raw_bytes = "caf\xC3\xA9";
my $text = decode('UTF-8', $raw_bytes, Encode::FB_CROAK);
die "unexpected text" unless $text eq 'café';
my $raw_again = encode('UTF-8', $text);
print length($raw_again), " bytes\n";
Encode::FB_CROAK makes malformed input fatal instead of silently replacing it. That is a useful boundary for configuration or identity data, where accepting damaged text can create ambiguity. Do not use low-level UTF-8 flag manipulation merely to make a warning disappear; it does not validate or convert the bytes.
6. Avoid the common security and debugging traps
Do not assume that visually similar Unicode strings are identical. Normalisation, confusable characters, bidirectional controls and malformed UTF-8 can affect comparisons, logs and access-control decisions. The perlunicode manual points to Unicode Technical Report 36 for the security implications. For identifiers, filenames and policy values, define an allowed format and compare the decoded data deliberately.
Keep byte and character operations separate. The bytes pragma is mainly a debugging tool, according to the manual. It can make ordinary string operations use byte semantics, which is exactly the wrong default for most text processing. Also remember that Unicode support does not automatically make every operating-system interface Unicode-aware; filenames, environment variables and command execution depend on the platform.
If an extension is not Unicode-aware, wrap it at one boundary: encode arguments to the format it documents, call it, then decode documented UTF-8 results. Check that module's manual first. Never guess an encoding from a display that happens to look correct.
Done means
- The program declares
use v5.12or later, and usesuse utf8when the source contains non-ASCII literals. - Text filehandles use an explicit
:encoding(UTF-8)layer. - Bytes entering from an external interface are decoded once, and bytes leaving are encoded once.
- Character limits use ordinary string operations; byte limits use an explicitly encoded value.
- Malformed or security-sensitive input is rejected or normalised according to a stated policy, not silently guessed.
cmp -- message.txt copy.txtreturns status 0 for the verified UTF-8 round trip.