Home / Alt manpages / perlunifaq(1)

  • perlunifaq(1)
  • User command
  • linux

Keep Perl Text and Bytes Sane with perlunifaq

You will finish with a small, repeatable rule for Perl text: decode bytes at the boundary, work with characters inside the program, and encode bytes at the boundary on the way out. You will also know which PerlIO layer to use and why a missing layer produces a misleading wide-character warning. Allow about fifteen minutes. You need Perl and the Encode module, both supplied by the installed perl-base and perl-doc packages on this machine.

1. Check the Perl version and available encodings

The local perlunifaq(1) manual is generated for Perl 5.38.2. It describes the current machine's Perl, not a guarantee that an older interpreter behaves identically. Check the interpreter before copying examples into a long-lived script:

$ perl -v
This is perl 5, version 38, subversion 2 (v5.38.2)

$ perl -MEncode -e 'print scalar(Encode->encodings(":all")), " encodings\n"'
180 encodings

The count can change when Perl or its modules change. The useful check is that Encode loads and that the encoding you need is available. To inspect names rather than a count:

$ perl -MEncode -le 'print for Encode->encodings(":all")' | grep -E '^(UTF-8|utf8|ISO-8859-1)$'
ISO-8859-1
UTF-8
utf8

Checkpoint: record the exact spelling of the external encoding. UTF-8 is the strict standard encoding name used for normal data exchange. Perl's utf8 name is more permissive and is not the default choice for validating incoming bytes.

2. Decode input before using it as text

An external file, socket, database or child process gives your program bytes. Convert those bytes to Perl characters immediately, while you still know which encoding they use. If the sender does not declare an encoding, investigate the protocol or document a deliberate guess. There is no reliable general-purpose encoding detector.

use strict;
use warnings;
use Encode qw(decode encode);

my $octets = "caf\xC3\xA9";
my $text   = decode('UTF-8', $octets);

print length($text), " characters\n";

That example prints 4 characters, because the decoded value contains c, a, f and e with an acute accent. The byte count and character count are different concepts. Keep them separate in variable names if that helps: $octets is a byte string and $text is character data.

Do not mix an undecoded UTF-8 byte string with a character string. Perl may implicitly treat the byte string as ISO-8859-1 and then upgrade it. That silently turns the bytes of a multibyte character into separate characters, a form of double encoding. A string that looks fine in a narrow test can therefore be damaged by one concatenation.

3. Encode output at the boundary

When text leaves the process, choose the receiver's encoding explicitly. This applies even when the receiver is another Perl program. For a UTF-8 file or pipe:

my $out_octets = encode('UTF-8', $text);
print $out_octets;

For standard output, make the boundary visible once rather than encoding every individual print:

use strict;
use warnings;
use open qw(:std :encoding(UTF-8));

print "Caf\x{E9}\n";

The use open form configures the standard handles for this process. It is useful when your program's standard input, output and error streams all have the same known encoding. If the streams differ, configure each handle separately instead. No elevated privileges are needed for these examples.

Without an output encoding, Perl can emit bytes for low code points and can emit UTF-8 for wider characters, while warning about a wide character when warnings are enabled. That produces an inconsistent stream. Treat the warning as a missing boundary decision, not as something to silence.

4. Let PerlIO do the conversion for a file

If every byte in a file uses one known encoding, attach an :encoding layer when opening it. The read layer decodes and the write layer encodes:

open my $in,  '<:encoding(UTF-8)', $input_path
    or die "open $input_path: $!";
open my $out, '>:encoding(UTF-8)', $output_path
    or die "open $output_path: $!";

while (my $line = <$in>) {
    print {$out} $line or die "write $output_path: $!";
}

close $out or die "close $output_path: $!";
close $in  or die "close $input_path: $!";

The input and output examples use UTF-8, but the layer can name another encoding supported by Encode. If the handle is already open, apply the same boundary with binmode $fh, ':encoding(UTF-8)'. Do not use :utf8 as a substitute for reading untrusted input: it accepts internal UTF-8 representations without the same validation and can create security problems with invalid byte sequences.

Checkpoint: after a conversion, verify the file as bytes, not just by opening it in an editor. For a known UTF-8 output:

$ perl -MEncode -0777 -e 'decode("UTF-8", <>); print "valid UTF-8\n"' output.txt
valid UTF-8

A decode failure is useful evidence that the file is not valid UTF-8, or that your assumption about its encoding is wrong. Do not solve that by replacing all bad bytes without first deciding what the sender meant.

5. Handle binary data and source files separately

Images and other binary formats are not text. A bare binmode $fh is enough to prevent newline translation on systems where that matters; do not decode arbitrary binary payloads. If text must be placed inside a binary stream, encode the text first and then combine the resulting bytes with the binary data.

The use utf8 pragma has a narrower job. Put it in a UTF-8 encoded Perl source file so Perl reads non-ASCII literals and identifiers correctly:

use utf8;
my $message = "Zażółć gęślą jaźń";

use utf8 does not decode input and does not encode output. It changes how the source file is read. Keep that distinction clear, especially when a script works from a terminal but fails when its output is redirected to a file.

6. Avoid the internal UTF8 flag as an application type

Perl's internal UTF8 flag describes an internal representation, not whether a value is text or binary. It can be absent for text stored in an eight-bit encoding and present after an internal conversion. Do not use utf8::is_utf8, _utf8_on or _utf8_off to decide what a value means. Track that meaning in your program's data flow or naming.

For character classes and case conversion on older code, the lexical unicode_strings feature can make Unicode rules explicit:

use feature 'unicode_strings';

my $text = "München";
print "has a word character\n" if $text =~ /\w+/;

The feature is available from Perl 5.14 and was partially available in 5.12. On this machine's Perl 5.38.2, it is also enabled by version feature bundles such as use v5.12. Keep it in the lexical scope where Unicode matching is required rather than relying on an interpreter-wide assumption.

Done means

  • You know the installed Perl version and have confirmed the required Encode name.
  • External bytes are decoded once before text operations.
  • Text is encoded once when it leaves the process, or a matching PerlIO layer does it.
  • Binary streams are kept separate from text and are not decoded by accident.
  • use utf8 is used only for UTF-8 source files, not for input or output.
  • You treat the UTF8 flag as an internal detail, not as a text-versus-binary label.