Home / Alt manpages / perlfaq6(1)

  • perlfaq6(1)
  • User command
  • linux

Practical Perl Regex Fixes from perlfaq6

You will finish with a small set of Perl regular-expression patterns that are easier to read, work across the right input boundaries, and do not accidentally treat literal user text as regex syntax. The examples target the installed Perl 5.38.2 and perlfaq6 version 5.20210520. Allow about 20 minutes, including time to try the checks against your own text.

You need a shell and the perl interpreter. The examples read standard input or literal strings and do not need elevated privileges. They do not edit files unless you redirect output yourself.

1. Check the Perl version and the input model

Start by confirming the interpreter that will run the examples:

$ command -v perl
/usr/bin/perl
$ perl -e 'print "$^V\n"'
v5.38.2

perlfaq6 is a FAQ, not a complete regex reference. Its central warning is that a pattern cannot match text that is not in the current string. Perl's default line-input loop gives you one line at a time, normally including its newline, so a pattern looking for text on two separate lines needs a different record size.

Checkpoint: decide whether your record is one line, a paragraph, or the whole file before choosing regex modifiers. This prevents the common distraction of changing /s and /m while the required text is not in memory at all.

2. Make a complicated pattern readable

The /x modifier ignores pattern whitespace and permits comments outside character classes. Paired delimiters also make substitutions containing slashes easier to inspect. This example removes simple HTML-like tags from one string:

$ perl -e '
my $text = q{<em>this</em> and <a path="/docs">that</a>};
$text =~ s{
    <                    # opening angle bracket
    (?:
        [^>\x27"]*        # unquoted content
        | " .*? "         # double-quoted content
        | \x27 .*? \x27   # single-quoted content
    )+
    >                    # closing angle bracket
}{}gsx;
print "$text\n";
'
this and that

This is still a deliberately limited example, not an HTML parser. When the input is XML or HTML, perlfaq6 recommends using a suitable module such as XML::LibXML or an HTML parser. A regex can remove a narrow, known form, but it does not understand nesting, quoted data in every context, or malformed markup.

3. Match across lines without guessing

Use /s when a dot should include newline characters. Use /m when ^ and $ should work beside embedded newlines. They solve different problems and can be combined.

To print text between START and END, read the whole input with -0777, then use a non-greedy capture and /s:

$ printf 'header\nSTART\nkeep this\nEND\ntrailer\n' \
  | perl -0777 -ne 'print "$1\n" while /START(.*?)END/sg'
keep this

The g modifier finds repeated regions. The ? makes the capture stop at the nearest END. Without it, a greedy .* can consume from the first START to the last END.

For line-oriented extraction, Perl's range operator is often clearer:

$ printf 'before\nSTART one\ninside\nEND one\nafter\n' \
  | perl -ne 'print if /START/ .. /END/'
START one
inside
END one

The range form prints complete lines, including both boundary lines. It is not a general nested-section parser. If START and END can nest, use a counter or a parser designed for the format.

4. Quote literal input before matching it

A variable interpolated into m// is still regex syntax. If the user enters P., the dot matches any character. Use \Q and \E, or the quotemeta function, when the value is meant to be literal:

$ perl -e '
my $text = q{Placido P. Octopus};
my $needle = q{P.};
print "regex match\n" if $text =~ /$needle/;
print "literal match\n" if $text =~ /\Q$needle\E/;
'
regex match
literal match

The first test succeeds because . is special. The second succeeds because Perl quotes the value before compiling the pattern. This is the safe default for a search box, filename filter, or other input where regex features were not explicitly requested.

If the user is intentionally supplying a regex, compile it with qr// inside an eval and handle invalid syntax. Otherwise a malformed pattern can terminate the program:

$ perl -e '
my $input = q{Unmatched ( paren};
my $pattern = eval { qr/$input/ };
if (defined $pattern) { print "valid pattern\n" }
else { print "invalid pattern: $@" }
'
invalid pattern: Unmatched ( in regex; marked by <-- HERE in m/Unmatched ( <-- HERE  paren/ at -e line 3.

5. Keep changing operations reversible

Perl's s/// changes the scalar in memory. A one-liner using -p prints the changed text, but it does not write the input file unless you add in-place editing. Treat -i as a destructive boundary: copy the file first, test the transformation on standard output, and only then choose an explicit backup suffix.

$ printf 'alpha\nBeta\n' | perl -pe 's/beta/gamma/gi'
alpha
gamma

For a file, preview first:

$ perl -pe 's/beta/gamma/gi' /path/to/input.txt > /tmp/input.txt.new
$ diff -u /path/to/input.txt /tmp/input.txt.new

If the diff is wrong, delete the temporary output and the original is untouched. If the diff is right, replace the original only after checking the destination and keeping a backup according to your normal change process. Do not use sudo merely because a regex is involved.

6. Use the FAQ as a boundary, not a pattern dump

perlfaq6 also covers locale-sensitive character classes, Unicode and multibyte text, repeated matching with /g, and performance traps such as $&. The practical rule is to identify the data model first. Use use locale when matching depends on the current locale, and use Perl's Unicode facilities when the input is text rather than arbitrary bytes.

Do not reach for a regex to parse balanced programming-language or markup structures. Backtracking regexes are not POSIX DFAs and do not promise worst-case performance for every pattern. Keep patterns bounded and specific when input is large or untrusted, and prefer a parser when the format has nesting or quoting rules.

Done means

  • You confirmed the Perl interpreter and version used for the test.
  • You chose a record size before changing /s or /m.
  • You used /x and comments when a pattern needed maintenance.
  • You quoted literal variable input with \Q...\E or quotemeta.
  • You compiled user-supplied regexes inside eval and handled errors.
  • You previewed substitutions and kept the original file before any replacement.