Home / Alt manpages / perlretut(1)

  • perlretut(1)
  • User command
  • linux

Build Reliable Perl Regex Checks and Extractors

You will finish with small Perl programs that test a whole value, extract fields from a line, replace text, and report failures without guessing what a pattern matched. The examples target the installed Perl 5.38.2 and its perlretut(1) tutorial. Allow 20 minutes if Perl is already installed.

Run the examples as an ordinary user. They only read literal input and print results. No root access, service restart or file change is needed.

1. Confirm the Perl version

Check the interpreter before relying on a feature. The local package is Perl 5.38.2, with perl-doc at the matching package version:

$ perl -e 'printf "%vd\n", $^V'
5.38.2

The tutorial's use re 'strict' pragma was introduced in Perl 5.22, so it is available here. Older interpreters may reject that pragma or apply different regexp fixes. Keep the version check in deployment notes when a script runs on more than one host.

Checkpoint

If the command prints a different major or minor version, test the script with that interpreter before putting it into a job.

2. Match a required shape

A match searches anywhere in the target by default. Anchors make the intended boundary visible. This example accepts exactly a lower-case word made from letters and digits:

$ perl -e 'use strict; use warnings; for my $value (qw(node-17 node17 xnode17)) { print "$value: ", $value =~ /^[a-z0-9]+$/ ? "valid\n" : "invalid\n" }'
node-17: invalid
node17: valid
xnode17: valid

^ anchors the start and $ anchors the end, while [a-z0-9]+ requires one or more permitted characters. Without the anchors, node17 would also match inside a longer value such as xnode17. That is a common validation error: finding an acceptable fragment is not the same as validating the complete input.

Matching is case-sensitive unless you add the /i modifier. For multiple lines in one string, /m changes the meaning of the line anchors; it does not split the string for you. Test that distinction explicitly before using a pattern against a whole file.

3. Extract named fields from a line

Parentheses capture text matched by a subpattern. Named captures are easier to review than positional values when a pattern grows:

$ perl -e 'use strict; use warnings; my $line = "user=ada id=42"; if ($line =~ /\Auser=(?<user>[a-z]+)\s+id=(?<id>\d+)\z/) { print "user=$+{user}\nid=$+{id}\n" } else { die "line has an unexpected format\n" }'
user=ada
id=42

\A and \z require the absolute start and end of the string. They avoid the special end-of-line behaviour associated with $, which can match before a final newline. \s+ accepts one or more whitespace characters, and \d+ accepts one or more digits. The captured values are available in $+{user} and $+{id} after a successful match.

If the input is external, treat a failed match as a normal parse failure. Do not use an undefined capture as if it proved that a field was absent: first check the whole match, then consume the named values.

4. Replace text without changing the original by accident

The s/// operator returns a changed string when used on a copy. Add /g when every occurrence should be replaced:

$ perl -e 'use strict; use warnings; my $text = "red, red, blue"; (my $copy = $text) =~ s/red/green/g; print "original: $text\nchanged:  $copy\n"'
original: red, red, blue
changed:  green, green, blue

Without /g, only the first match is replaced. Writing the copy first makes that state change obvious. If you intend to alter the original variable, write $text =~ s/red/green/g and verify the result immediately.

Choose delimiters that keep a pattern readable. For a path, m{/var/log} avoids a row of escaped slashes. A metacharacter such as +, . or ? has special meaning, so escape it when you need the literal character. A value interpolated into a pattern is not automatically literal; quote or constrain external values before using them as a pattern.

5. Make dense patterns reviewable

Use the /x modifier for a pattern that needs explanation. It permits layout whitespace and comments, but unescaped spaces in the pattern no longer match spaces in the input:

$ perl -e 'use strict; use warnings; my $value = "ticket=AB-123"; if ($value =~ /^ ticket= (?<ticket> [A-Z]{2}-\d{3} ) $/x) { print "$+{ticket}\n" }'
AB-123

In a longer program, put that pattern in a qr// variable so it can be named and reused. Add use re "strict" while developing patterns. Perl can then flag constructs that are legal but likely to express a mistake:

$ perl -e 'use strict; use warnings; no warnings "experimental::re_strict"; use re "strict"; my $pattern = qr/^ticket=[A-Z]{2}-\d{3}$/; print "match\n" if "ticket=AB-123" =~ $pattern'
match

Do not treat a successful compile as proof that the business rule is correct. Keep a matching and non-matching test beside the pattern, especially around optional fields and repeated groups.

6. Diagnose a surprising result

Start with the smallest test string and print the result of the operation you actually care about. The regexp normally finds the earliest possible match, and repetition is greedy, so .* can consume more than expected before backtracking. Prefer a narrow character class when the format gives you one:

$ perl -e 'use strict; use warnings; my $line = "name=Ada;role=admin"; if ($line =~ /\Aname=([^;]+);role=([^;]+)\z/) { print "name=$1 role=$2\n" }'
name=Ada role=admin

Here [^;]+ stops at the delimiter. A pattern such as name=(.*);role=(.*) is less precise and can become expensive on large, unexpected input. If you need to see how a pattern is compiled and executed, temporarily add use re 'debug'; remove it after diagnosis because it writes verbose tracing to standard error.

When the pattern is built from a user-supplied value, do not allow that value to become arbitrary regexp syntax. Use a literal-string approach such as quotemeta where appropriate, and set a sensible input size limit for data you do not control.

Done means

  • The script's Perl version has been checked and any version-specific pragma is deliberate.
  • Whole-value checks use explicit anchors and distinguish a failed match from a missing capture.
  • Named captures extract fields, and substitutions use /g only when every occurrence should change.
  • Patterns containing external text do not treat that text as trusted regexp syntax.
  • A narrow test case, a failing case and the expected output are recorded for each important pattern.