Write Perl That Survives ASCII and EBCDIC
You will finish with a small set of Perl patterns that behave predictably on ASCII and EBCDIC systems: Unicode escapes for character identity, explicit encodings for I/O, portable regular expressions, and a deliberate sort order. The installed reference is perlebcdic(1) from Perl v5.38.2, although the local machine is an ASCII Linux host and cannot execute EBCDIC-specific behaviour.
The route
Jump straight to the step you need, or tick off Done means at the end.
Allow 25 minutes. You need Perl, the standard Encode module and a shell. No root access is required. The examples create files only in a temporary directory and remove them at the end; do not adapt the cleanup command to a directory containing real data.
1. Check the Perl version and character model
Start by recording the interpreter you are testing. This is an ordinary, read-only check:
$ perl -v | sed -n '1,4p'
This is perl 5, version 38, subversion 2 (v5.38.2)
(...) built for x86_64-linux-gnu-thread-multi
The exact build string will differ. Perl's documentation says the core worked on z/OS again from v5.22, while support on other EBCDIC systems is less certain. Treat the target operating system, code page and Perl build as part of your test matrix, not as interchangeable details.
Checkpoint: write down the target's EBCDIC code page. The documentation discusses CCSID 0037, CCSID 1047 and POSIX-BC, and says Perl recognises those three commonly used sets. An EBCDIC byte value is not automatically an ASCII or Unicode code point: for example, the letter A is commonly 193 in EBCDIC but 65 in Unicode.
2. Replace byte-number assumptions with character names
Escapes such as \n and \t are useful for ordinary controls. Fixed byte escapes are different. The manpage gives \xC1 as an EBCDIC spelling for A, but that value is not portable to an ASCII machine. Use a Unicode code point or character name when the character is the thing you mean:
use strict;
use warnings;
use feature 'unicode_strings';
my $letter = "\N{U+0041}";
my $y_diaeresis = "\N{U+00FF}";
printf "letter=%s, code point=U+%04X\n", $letter, ord($letter);
printf "named character=%s\n", $y_diaeresis;
On an ASCII build, ord reports the native value used by that build. The useful property here is the character identity, not a guessed byte number. This form requires Perl 5.22 or later for the portable character-class examples in the next step; the installed interpreter is newer.
Do not replace every byte operation with a character operation. A protocol, checksum or file format may deliberately specify bytes. First decide whether a value is text or binary data, then keep that decision visible in the code.
3. Make character classes portable
A range such as [\x00-\x1F] describes ASCII byte values. It does not reliably describe the same characters on EBCDIC. For Perl 5.22 and later, use Unicode escapes when you need the ASCII or Latin-1 character ranges:
sub is_print_ascii {
my ($char) = @_;
return $char =~ /[\N{U+0020}-\N{U+007E}]/;
}
sub is_delete {
my ($char) = @_;
return $char eq "\N{U+007F}";
}
printf "%d %d\n", is_print_ascii('A'), is_delete("\N{U+007F}");
Expected output on a normal Perl build is:
1 1
For a broader test, the manpage suggests character properties such as [[:print:]] and [[:cntrl:]], with the /a modifier when an ASCII-only interpretation is required. Be precise about the requirement: "printable in ASCII" and "printable in this locale" are different tests.
4. Put encoding at the I/O boundary
Keep your program's text as characters, then choose an encoding when opening each external channel. The manpage recommends PerlIO layers and shows separate ASCII, Latin-1, UTF-8 and EBCDIC outputs. This small example writes Latin-1 explicitly:
use strict;
use warnings;
open my $out, '>:encoding(latin1)', 'message.txt'
or die "message.txt: $!";
print {$out} "caf\N{U+00E9}\n";
close $out or die "message.txt: $!";
The filename is relative to the current directory. Use a known temporary directory while experimenting. Never assume that a filesystem being called "ASCII" or "EBCDIC" changes what the PerlIO layer does; the documentation says the layer uses raw I/O internally and applies the selected channel encoding.
When translating an in-memory string rather than opening a file, Encode::from_to can convert between an EBCDIC code page and Latin-1:
use Encode qw(from_to);
my $text = "hello";
from_to($text, 'latin1', 'cp1047');
from_to($text, 'cp1047', 'latin1');
print "$text\n";
Do not guess the code page. cp37, cp1047 and posix-bc are not interchangeable, particularly for variant punctuation characters. Verify the source system's CCSID before selecting a name.
5. Make sorting an explicit policy
Native sorting produces different results. In ASCII, digits sort before uppercase letters, then underscore, then lowercase letters. In EBCDIC, underscore, lowercase letters, uppercase letters and digits have a different order. This can change reports, tests and generated file lists:
my @names = qw(Dr. dr.);
print join(', ', sort @names), "\n";
If native order is the requirement, leave the sort alone and test expected results on each target. If output must have one cross-platform order, transform each comparison to Unicode order:
sub native_to_unicode {
my ($string) = @_;
return $string if ord('A') == 65;
my $output = '';
for my $i (0 .. length($string) - 1) {
$output .= chr(utf8::native_to_unicode(ord(substr($string, $i, 1))));
}
utf8::upgrade($output) if utf8::is_utf8($string);
return $output;
}
sub unicode_order { native_to_unicode($a) cmp native_to_unicode($b) }
print join(', ', sort unicode_order @names), "\n";
This helper costs more than a native comparison. Use it where stable interchange or presentation order matters, not automatically in every hot loop. Also remember that Perl deliberately randomises hash order on both families of platform, so sorting keys is a separate step whenever deterministic output matters.
6. Check boundaries before shipping
Network formats are another common trap. The manpage says most socket programming assumes ASCII character encodings in network byte order, while host web servers may translate EBCDIC output. Do not rely on that translation for an application protocol. Define the wire encoding, encode before sending, decode after receiving, and test both ends.
Quoted-printable and URL percent-encoding are especially easy to get wrong if the hexadecimal number is treated as a native character value. Convert through Unicode with utf8::native_to_unicode when producing an ASCII-defined representation, and use utf8::unicode_to_native when decoding it back to a native character. For a protocol library, prefer its tested encoder rather than copying a partial substitution.
Checkpoint: run the test suite on at least one ASCII build and one EBCDIC build if portability is a release requirement. Compare semantic text, explicitly encoded bytes, sorted output and checksums separately. An EBCDIC checksum can legitimately differ from the checksum of the same content translated to ASCII.
Done means
- You recorded the Perl version and target EBCDIC code page.
- Character identity uses Unicode escapes or names instead of guessed byte values.
- Character classes say whether they mean Unicode, ASCII or locale properties.
- Each I/O channel declares its encoding, and the code page is verified.
- Sorting and network representations have an explicit, tested policy.
- You kept byte-level formats separate from text processing.