Make Perl Unicode Text Predictable with Explicit I/O
You will finish with a small Perl program that reads UTF-8 as characters, processes it without confusing bytes and text, and writes valid UTF-8. The examples target the installed Perl 5.38.2 from the perl 5.38.2-3.2ubuntu0.6 package and use the perluniintro guidance shipped by perl-doc.
The route
Jump straight to the step you need, or tick off Done means at the end.
- 1. Check the Perl version and make a scratch directory
- 2. Create Unicode text by code point
- 3. Decode input at the file boundary
- 4. Encode output explicitly
- 5. Use Encode when you have bytes, not a file handle
- 6. Keep character and byte measurements separate
- 7. Avoid the traps that look like successful output
Allow about 20 minutes. You need a shell, Perl and a temporary working directory. The examples create or replace only files beneath /tmp/perl-unicode-demo. Do not point the output example at a valuable file until you have a backup: opening a path with > truncates it.
1. Check the Perl version and make a scratch directory
Start by confirming which interpreter will run the program:
$ perl -v | sed -n '1,6p'
This is perl 5, version 38, subversion 2 (v5.38.2) built for x86_64-linux-gnu-thread-multi
Create an isolated directory for the test file:
$ mkdir -p /tmp/perl-unicode-demo
$ cd /tmp/perl-unicode-demo
Checkpoint: if perl -v reports an older interpreter, keep the explicit I/O layers in the examples, but check the version-specific caveats in perluniintro before relying on newer syntax or Unicode fixes.
2. Create Unicode text by code point
Perl can create a character from its Unicode code point with chr. Hexadecimal is the usual notation in Unicode documentation. Use a character name with \N{...} when that makes the source clearer:
use strict;
use warnings;
use feature 'unicode_strings';
my $text = "caf\N{LATIN SMALL LETTER E WITH ACUTE}";
my $smiley = chr(0x263A);
binmode STDOUT, ':encoding(UTF-8)' or die "cannot set stdout: $!";
print "$text $smiley\n";
Save that as create.pl and run it:
$ perl create.pl
café ☺
use feature 'unicode_strings' gives the script the Unicode string behaviour described by the manual. On Perl 5.38 it is also selected by use v5.12 or higher. This is separate from use utf8: use utf8 is needed when the program source itself contains UTF-8 characters, whereas \N{...} keeps this example ASCII.
3. Decode input at the file boundary
A file containing UTF-8 bytes does not become Unicode text merely because it contains UTF-8. Declare the encoding while opening it, so Perl decodes bytes as they are read:
use strict;
use warnings;
use feature 'unicode_strings';
binmode STDOUT, ':encoding(UTF-8)' or die "cannot set stdout: $!";
open my $in, '<:encoding(UTF-8)', 'input.txt'
or die "cannot open input.txt: $!";
while (my $line = <$in>) {
chomp $line;
print length($line), " characters: $line\n";
}
close $in or die "cannot close input.txt: $!";
Make a known UTF-8 input file and run the program:
$ printf '%s\n' 'café' 'Å' > input.txt
$ perl read.pl
4 characters: café
2 characters: Å
The second line has an A followed by a combining acute accent. Perl counts code points, so length returns 2 even though many readers see one displayed grapheme. A regular expression using \X can match an extended grapheme cluster when that distinction matters.
Do not use < without an encoding layer for text whose encoding you know. That leaves the bytes uninterpreted and pushes the mistake into later processing.
4. Encode output explicitly
Use the matching output layer when writing Unicode text. :encoding(UTF-8) converts Perl characters to UTF-8 bytes and validates the conversion rules:
use strict;
use warnings;
use feature 'unicode_strings';
open my $out, '>:encoding(UTF-8)', 'output.txt'
or die "cannot open output.txt: $!";
print {$out} "caf\N{LATIN SMALL LETTER E WITH ACUTE} ", chr(0x263A), "\n";
close $out or die "cannot close output.txt: $!";
Verify the result without relying on how your terminal renders it:
$ perl write.pl
$ od -An -tx1 -c output.txt
63 61 66 c3 a9 20 e2 98 ba 0a
c a f 303 251 342 230 272 \n
The accented é is two UTF-8 bytes and the smiley is three. If you print a character above code point 0xFF to a stream with no PerlIO encoding layer, Perl can emit its internal bytes and warn about a wide character. That warning is a symptom of an unspecified output boundary, not something to silence blindly.
For an already-open stream such as standard output, set the layer with binmode:
binmode STDOUT, ':encoding(UTF-8)' or die "cannot set stdout: $!";
print "café\n";
5. Use Encode when you have bytes, not a file handle
Sometimes a library or socket gives you a scalar containing raw bytes. Decode those bytes once, at the point where you know their encoding:
use strict;
use warnings;
use Encode qw(decode FB_CROAK);
my $bytes = "caf\xC3\xA9\n";
my $text = decode('UTF-8', $bytes, FB_CROAK);
print "$text";
FB_CROAK makes malformed UTF-8 an error instead of quietly replacing or dropping data. For file handles, the :encoding(UTF-8) layer is usually simpler because it performs the conversion while reading and writing.
Checkpoint: an input boundary should perform one decode, and an output boundary should perform one encode. Applying UTF-8 encoding to data that is already UTF-8 bytes is the classic double-encoding bug. The resulting file may appear as sequences such as café after a second pass.
6. Keep character and byte measurements separate
Normal Perl string operations work on Unicode characters once the scalar is Unicode. If a protocol or binary format needs byte counts, request that explicitly with the bytes pragma:
use strict;
use warnings;
my $text = chr(0x100);
print length($text), " character\n";
{
use bytes;
print length($text), " UTF-8 bytes\n";
}
$ perl lengths.pl
1 character
2 UTF-8 bytes
Do not use use bytes as a general Unicode switch. It changes operations such as length to byte semantics for the scope where it is active. Likewise, utf8::is_utf8 reports Perl's internal UTF-8 flag; it does not prove that a scalar contains valid UTF-8 or that it came from a UTF-8 file.
7. Avoid the traps that look like successful output
- Source encoding: if the source file itself contains literal non-ASCII characters, add
use utf8or write the literals with\N{...}. This concerns the program source, not input files. - Normalisation: a precomposed character and a base character plus combining mark can compare differently. Perl's default string comparison uses code points. Use
Unicode::Normalizewhen your application needs canonical equivalence. - Regular expressions:
[A-Za-z]is not a general Unicode alphabet test. Use Unicode properties such as\p{Alpha}when that is the rule you actually need. - Binary handles: do not combine
sysreadorsyswritewith character encoding layers. The local manual says those operations behave badly there and have been deprecated in that use since Perl 5.24. - Terminal settings: a correct UTF-8 file can still display incorrectly if the terminal or locale expects another encoding. Inspect bytes with
odbefore changing the program.
Done means
- The script's Perl version is known, here
5.38.2. - Known text input is opened with
<:encoding(UTF-8). - Text output uses
>:encoding(UTF-8)or an explicitbinmodelayer. - Raw byte strings are decoded once with a named encoding and a deliberate error policy.
- Character lengths and byte lengths are measured with the intended semantics.
- No file is encoded twice, and the scratch files can be removed safely with
rm -rf /tmp/perl-unicode-demowhen you have finished checking them.