Home / Alt manpages / perlpacktut(1)

  • perlpacktut(1)
  • User command
  • linux

Use Perl pack and unpack for Fixed-Width and Binary Data

You will finish with a small Perl toolkit for turning values into byte strings, reading fixed-width records, and checking binary data without guessing what each byte means. The examples match the perlpacktut(1) tutorial installed with Perl 5.38.2 on this machine. Allow about 20 minutes if you can run Perl locally.

You need Perl and a shell. No elevated privileges are required. Work in a scratch directory if you are testing input from another system. pack and unpack do not edit files by themselves, but a script that writes packed data can still overwrite a destination, so use a new output path while learning.

1. Confirm the Perl version and make a round trip

Check the interpreter first. This matters because the installed tutorial is versioned, and packing rules such as native integer sizes can vary between platforms:

$ perl -v | sed -n '1,4p'
This is perl 5, version 38, subversion 2 (v5.38.2)

The first useful model is simple: pack converts values into a byte sequence using a template; unpack reads values from bytes using a template. A hex string is an easy way to see the result:

use strict;
use warnings;

my $bytes = pack 'C*', 65, 66, 67;
my $hex = unpack 'H*', $bytes;
print "$hex\n";

my @values = unpack 'C*', $bytes;
print join(',', @values), "\n";

Run it from a file or with a here document. The expected output is:

414243
65,66,67

C* means an unsigned byte for every supplied value when packing, or every remaining byte when unpacking. H* displays the bytes as hexadecimal digits. The star is a repeat count meaning "use the remainder", not a literal character in the data.

Checkpoint

If the first line is not 414243, check that you ran the intended Perl interpreter and that the template is quoted exactly.

2. Read a fixed-width text record

Fixed-width data is where unpack often removes the most fragile counting. Suppose each input line contains a ten-character date, one skipped separator, a 27-character description, another separator, and a seven-character amount:

use strict;
use warnings;

my $line = "01/24/2001 Zed's Camel Emporium          1147.99\n";
chomp $line;
my ($date, $description, $amount) = unpack 'A10xA27xA7', $line;

for ($date, $description, $amount) {
    s/\s+\z//;
}
print "date=[$date]\ndescription=[$description]\namount=[$amount]\n";

The output should be:

date=[01/24/2001]
description=[Zed's Camel Emporium]
amount=[1147.99]

A reads a text field and pads or truncates it to the width after the code. x skips bytes while unpacking and inserts zero bytes while packing. The spaces inside an A field are part of the field, which is why the example trims trailing whitespace after unpacking.

Do not use this template for a delimiter-separated format merely because one sample line happens to fit. If a description can contain a variable number of characters, use the format's delimiter or parser instead. Also remember that these widths are byte-oriented; a non-ASCII character may occupy more than one byte in UTF-8.

3. Pack integers for a defined byte order

Numeric templates need a deliberate choice of size, signedness, and byte order. The portable network-order codes are useful when a protocol specifies big-endian integers. This example packs two unsigned 16-bit values and immediately reads them back:

use strict;
use warnings;

my $packet = pack 'n2', 4660, 258;
print unpack('H*', $packet), "\n";
my ($first, $second) = unpack 'n2', $packet;
print "$first $second\n";
12340102
4660 258

n is an unsigned short in network, or big-endian, byte order. The corresponding little-endian code is v. Native codes such as s, i, and I describe the host Perl was built for, so use them only when the format is explicitly host-native. A template such as i does not promise a fixed number of bytes across machines.

Safety boundary

Do not send packed values to a network peer until the peer's specification has fixed the width, byte order, and signedness. A round trip on your own host only proves that your two templates agree with each other.

4. Inspect an unknown byte string

When a file or socket gives you bytes but no trustworthy description, inspect first and interpret second. This script prints two hexadecimal digits per byte:

use strict;
use warnings;

my $data = "AB\x00\xff";
my @hex = unpack 'H2' x length($data), $data;
print join(' ', @hex), "\n";
41 42 00 ff

This output tells you what arrived, not what it means. Avoid treating arbitrary bytes as text or assuming ASCII when the format has not said so. For Unicode text, the tutorial also documents the U template for Perl's character values. In ordinary application code, keep the distinction clear: use a documented character encoding such as UTF-8 and decode it with the appropriate Perl encoding tools rather than using a numeric template as a general text decoder.

5. Check the result and recover from mistakes

For a new template, test both directions with known values and inspect the byte length:

my $bytes = pack 'n2', 4660, 258;
die "wrong length\n" unless length($bytes) == 4;
die "round trip failed\n" unless join(',', unpack('n2', $bytes)) eq '4660,258';
print "template verified\n";

If a fixed-width result is shifted, count the format's separators and widths again. If a number looks implausible, compare the documented byte order and signedness before changing the data. If the destination file was accidentally replaced by a script, stop using it and restore it from your normal backup or source copy; pack has no undo operation.

Done means

  • You can explain which side of the conversion uses pack and which uses unpack.
  • Your fixed-width template names every field width and skipped byte.
  • Your binary template states its integer size, signedness, and byte order.
  • You have checked a hex representation and a round-trip value before using real data.