Handle Unicode Text Safely in Perl with perlunicook
You will finish with a small Perl program that reads UTF-8 text, treats it as characters internally, writes UTF-8 deliberately, and counts user-visible characters without splitting combining marks. The examples follow perlunicook(1) as installed by perl-doc 5.38.2 on this machine.
The route
Jump straight to the step you need, or tick off Done means at the end.
Allow about 15 minutes. You need Perl 5.36 or later for the standard preamble used here. The core examples need no elevated privileges and do not alter system files. A few optional cookbook recipes use CPAN modules; they are clearly marked and are not needed for this workflow.
1. Start with a Unicode-aware program
Save this as /tmp/unicode-check.pl. The preamble establishes three separate facts: the source file is UTF-8, Perl warnings are strict about encoding faults, and standard streams use UTF-8.
#!/usr/bin/env perl
use v5.36;
use utf8;
use warnings qw(FATAL utf8);
use open qw(:std :encoding(UTF-8));
use charnames qw(:full :short);
my $text = "Crème brûlée: café\n";
print $text;
Run it with the system Perl:
$ perl /tmp/unicode-check.pl
Crème brûlée: café
Checkpoint: if the accents print as replacement characters or a warning is raised, check the terminal and locale before changing the program. The preamble controls Perl's streams, not a terminal emulator that cannot display UTF-8.
2. Keep bytes at the boundary
Inside Perl, work with decoded characters. Decode bytes when you know their encoding, and encode only when producing bytes for a file, socket, or other external interface. For a file whose encoding is known, attach the encoding to the handle rather than scattering calls to encode and decode through the processing loop.
use v5.36;
use utf8;
use warnings qw(FATAL utf8);
use open qw(:std :encoding(UTF-8));
open my $in, '<:encoding(UTF-8)', '/path/to/input.txt'
or die "open input: $!";
open my $out, '>:encoding(UTF-8)', '/path/to/output.txt'
or die "open output: $!";
while (my $line = <$in>) {
print {$out} $line;
}
close $in or die "close input: $!";
close $out or die "close output: $!";
The angle brackets are escaped because this is HTML, but the Perl operators are ordinary < and > in the file. The same pattern works for another known encoding, such as :encoding(UTF-16) or :encoding(cp1252). Do not guess an encoding from a successful open call: a wrong encoding can still produce plausible but corrupted text.
For a one-off byte string, use the three-argument forms so malformed input is fatal:
use Encode qw(encode decode);
my $characters = decode('UTF-8', $bytes, 1);
my $new_bytes = encode('UTF-8', $characters, 1);
Do not use encode or decode as a substitute for setting a stream layer. The manpage's normal rule is to set the file encoding when opening the file or with binmode.
3. Normalise at the application boundary
The same visible text may have one precomposed codepoint or a base character followed by a combining mark. If you compare, index, or search text from different sources, choose a normalisation policy. The cookbook's general pattern is NFD on input and NFC on output.
use Unicode::Normalize qw(NFD NFC);
my $internal = NFD($characters);
# compare or search $internal here
my $display = NFC($internal);
print {$out} $display;
Use NFKC or NFKD only when compatibility equivalence is what the application wants, such as improving search recall. Compatibility normalisation can change distinctions that matter to identifiers or display text.
4. Count and slice graphemes, not codepoints
A user-visible character is not always one codepoint. The regex escape \X matches a grapheme cluster, so it is the right primitive for display-oriented limits and extraction.
my $label = "brûlée";
my $count = 0;
$count++ while $label =~ /\X/g;
my ($first_five) = $label =~ /^(\X{5})/;
say "graphemes=$count";
say "prefix=$first_five";
Expected output is:
graphemes=6
prefix=brûlé
For repeated grapheme operations, the optional Unicode::GCString module provides substr, length, and display-column calculations. Ordinary substr, length, and reversal operate at a different level and can separate a combining mark from its base character.
5. Make matching and sorting explicit
Perl's Unicode-aware regular expressions provide properties such as \p{Latin}, \p{Greek}, and \p{alpha}. Use them when the rule is about Unicode character properties. If a protocol really requires ASCII character classes, use the /a regex modifier for that expression, rather than assuming every \d match is an ASCII digit.
my @words = grep { /\A\p{Latin}+\z/ } qw(cafe café 東京);
say join ',', @words;
Unicode case folding is available through fc when you need case-insensitive comparisons or sorting:
use feature 'fc';
my @sorted = sort { fc($a) cmp fc($b) } @words;
Do not use plain cmp as an alphabetic collation rule. For user-facing ordering, Unicode::Collate implements the Unicode Collation Algorithm, and Unicode::Collate::Locale supplies locale-specific rules. A locale is a policy choice, not a universal improvement.
6. Verify the boundary you changed
Check the program's syntax, then inspect the output bytes. These commands are read-only apart from the temporary files you created:
$ perl -c /tmp/unicode-check.pl
/tmp/unicode-check.pl syntax OK
$ perl /tmp/unicode-check.pl | od -An -tx1
43 72 c3 a8 6d 65 20 62 72 c3 bb 6c c3 a9 65 3a
20 63 61 66 c3 a9 0a
The bytes c3 a8 represent è, and c3 a9 represents é in UTF-8. If you need to remove the temporary test, use rm -- /tmp/unicode-check.pl after checking its exact path. That deletion is irreversible, so preserve it if you want a reproducible test case.
Common traps
- Missing
use utf8: UTF-8 literals and identifiers in the source are not declared correctly. - Confusing UTF-8 with Unicode: UTF-8 is an external byte encoding; Perl's internal text and a file's byte stream are different concerns.
- Using
PERL_UNICODEcasually:-CS,-CD, and-CSDAaffect different stream and argument defaults. Prefer a local, visibleuse opendeclaration when the policy belongs to one program. - Assuming a character is one codepoint: use
\Xfor user-visible limits, extraction, reversal, and counting. - Ignoring warnings: keep
warnings qw(FATAL utf8)while finding an encoding boundary. Silencing a warning can turn bad input into silent data loss.
Done means
- The source declares
use utf8and the standard UTF-8 stream layers. - Known input and output encodings are attached to file handles.
- Normalisation is a deliberate application decision.
- Display-oriented counts and slices use grapheme clusters.
- Matching and sorting rules state whether they are Unicode, ASCII, or locale-specific.
perl -cpasses and a byte-level check confirms the expected UTF-8 output.