Perl Regexes in Practice: Match, Extract and Replace Safely
You will build small Perl regular expressions that find text, extract fields, process every match, split input and perform controlled replacements. The examples target Perl v5.38.2, the version installed with the local perl-doc package. Allow about 20 minutes. You need a shell and Perl; no elevated privileges are required, and every example works on data held in memory.
The route
Jump straight to the step you need, or tick off Done means at the end.
Save the examples in a temporary file if you want to edit them. The commands below use perl -e so each result is easy to copy and verify. A regular expression normally searches anywhere in its target string, is case-sensitive, and stops at the earliest position where it can match. Those three defaults explain many surprising results.
1. Make one match and check its status
Use the match operator with =~ to associate a string with a pattern:
$ perl -e 'print "matched\n" if "Hello World" =~ /World/;'
matched
The expression is true because World occurs in the target. The pattern is not implicitly anchored, so it does not need to consume the whole string. Perl is case-sensitive here: /world/ does not match Hello World. Use !~ when the useful branch is the non-match.
Checkpoint: compare a substring with a whole-string check:
$ perl -e 'print "partial\n" if "housekeeper" =~ /keeper/; print "whole\n" if "housekeeper" =~ /^housekeeper$/;'
partial
whole
The anchors ^ and $ constrain the beginning and end. The documented $ anchor can also match just before a final newline, so remove or account for line endings when validating records.
2. Match alternatives and character shapes
Put alternatives on either side of |, and group them when they belong to a larger pattern:
$ perl -e 'print "found\n" if "cats and dogs" =~ /(cat|dog)/;'
found
At one position Perl tries alternatives from left to right, but it searches the string for the earliest position first. In cats and dogs, cat wins even if dog is listed first, because cat starts earlier.
Character classes describe one character from a set. Ranges such as [0-9] and shorthand classes such as \d, \s and \w are convenient, but their Unicode behaviour can be broader than ASCII unless you select the appropriate mode. For a simple time-shaped field:
$ perl -e 'print "time-shaped\n" if "09:42:17" =~ /^\d\d:\d\d:\d\d$/;'
time-shaped
Escape metacharacters when you mean literal punctuation. For example, 2+2 needs /2\+2/; an unescaped + is a repetition operator. If a pattern contains confusing or nested syntax, add use re "strict" in a script so Perl can report constructs that are legal but may not mean what you intended.
3. Capture fields and repeat a search
Parentheses both group a pattern and capture the text matched inside them. The captures are available as $1, $2 and so on after a successful match:
$ perl -e '$time = "09:42:17"; if ($time =~ /^(\d\d):(\d\d):(\d\d)$/) { print "hour=$1 minute=$2 second=$3\n"; }'
hour=09 minute=42 second=17
Capture numbering follows opening-parenthesis order. Keep captures that are only for grouping out of the numbering scheme in more complex patterns, or carefully update every later reference when the pattern changes. Inside a pattern, use backreferences such as \g1, not $1, to require the same captured text again:
$ perl -e 'print "doubled\n" if "the the" =~ /^(\w+)\s+\g1$/;'
doubled
To find every word, add /g and use a loop. In scalar context Perl retains the current position between successful matches:
$ perl -e '$x = "cat dog house"; while ($x =~ /(\w+)/g) { print "$1 at ", pos($x), "\n"; }'
cat at 3
dog at 7
house at 13
Changing the target or allowing a match to fail resets that position. Add /c to /gc when a failed attempt should leave the current position intact. In list context, /g instead returns all matches, or all captured values when the pattern has groups.
Checkpoint: decide which result you need before choosing context. A loop gives you positions and per-match logic; a list assignment gives you a compact collection.
4. Replace deliberately
Use s/// for substitution. The first match is replaced by default; add /g for every occurrence:
$ perl -e '$x = "I batted 4 for 4"; $x =~ s/4/four/g; print "$x\n";'
I batted four for four
Replacement text is processed like a double-quoted string, and captures can be reused. The non-destructive /r modifier returns a changed copy while preserving the original:
$ perl -e '$x = "I like dogs."; $y = $x =~ s/dogs/cats/r; print "$x\n$y\n";'
I like dogs.
I like cats.
Security checkpoint
Treat /e as executable Perl, not as a formatting switch. It evaluates the replacement expression and can run arbitrary code if untrusted input reaches it. Do not use it on input you have not validated; prefer a normal replacement or explicit code that checks allowed values first.
5. Split input and protect output files
split /pattern/, string returns the pieces separated by the pattern. Use \s+ to tolerate more than one space:
$ perl -e '@word = split /\s+/, "Calvin and Hobbes"; print join("|", @word), "\n";'
Calvin|and|Hobbes
A capture in the separator is included in the returned list. That is useful when the delimiter itself matters, but it can also create unexpected extra fields. Check the list with join before feeding it to later code.
No example here changes a file. If you adapt s/// to rewrite one, write to a new destination first, verify it, then replace the old file only after the check passes. Shell redirection with > truncates an existing destination before Perl runs, so use a deliberately new name or a temporary file in the same directory. Recovery is then simple: discard the temporary output and keep the original.
6. Diagnose the usual false matches
- A lowercase pattern does not match uppercase text unless you add the
/imodifier, as in/yes/i. - A pattern without
^and$can accept a valid-looking fragment inside a larger invalid string. .*is greedy. It consumes as much as possible while still allowing the rest of the pattern to succeed. Prefer a narrower character class or a lazy quantifier when the boundary is known.- Inside a character class,
-,],^in first position and\have special meanings. Escape them or place them where they are literal. - Remember that
$1is for code after the match, while\g1is for a backreference inside the pattern.
When a pattern becomes hard to review, use a named script, add use strict; and use warnings;, and test both accepted and rejected inputs. Keep the input examples small enough that you can see which character was matched.
Done means
- You can distinguish an anywhere match from a whole-string match.
- You can use classes, alternatives, anchors and escapes without treating punctuation as ordinary text.
- You can capture fields, choose scalar or list context, and understand how
/gadvances. - You use
s///gfor repeated replacements ands///rwhen the original must remain unchanged. - You treat
/eas executable code and protect any file replacement with a verified temporary output.