Home / Alt manpages / perlrebackslash(1)

  • perlrebackslash(1)
  • User command
  • linux

Perl regular-expression backslashes: escapes that stay predictable

You will write and test Perl regular expressions that use backslashes for literal punctuation, character classes, Unicode characters, anchors and backreferences. Allow about 15 minutes for the examples. No elevated privileges are needed: this guide only runs short, read-only Perl one-liners in your shell.

The examples were checked with Perl 5.38.2 from Ubuntu package perl 5.38.2-3.2ubuntu0.6. The installed manual is dated 14 September 2026. Perl's regular-expression syntax has version-specific additions, so treat a newer or much older interpreter as a separate compatibility target.

1. Confirm the interpreter before reading a result

First check which Perl will parse the pattern. This prevents a shell path or an unexpected interpreter from becoming the real source of a failure.

$ command -v perl
/usr/bin/perl
$ perl -e 'printf "%vd\n", $^V'
v5.38.2

Checkpoint: the version should be the one you intend to support. The rest of this guide uses normal, unprivileged commands and does not edit files or configuration.

2. Decide whether the backslash is quoting or an escape

Perl gives a backslash one of two jobs. Before an ASCII punctuation character it removes that character's pattern meaning, so \| matches a literal vertical bar. Before an ASCII letter or digit it may start a defined sequence such as \d or \x{...}. A backslash before an unused letter can warn today and gain meaning in a future Perl release, so warnings are useful.

The backslash itself must be escaped when it is the character to match. In a single-quoted shell argument, this pattern has two backslashes and matches one:

$ perl -e 'print "matched\n" if "path\\name" =~ /\\/'
matched

Do not confuse Perl's pattern with the shell's quoting. Single quotes stop the shell interpreting most backslashes; double quotes can interpret some shell sequences before Perl sees them.

3. Use fixed-character and Unicode escapes for data

Use the named fixed escapes \t, \n, \r, \f, \a and \e when the character is the point of the pattern. \n is Perl's logical newline. For a platform-independent control character or Unicode code point, \N{...} is clearer than relying on a numeric character value.

$ perl -Mutf8 -CS -e 'print "tab\n" if "a\tb" =~ /\t/; print "snowman\n" if "☃" =~ /\x{2603}/; print "A\n" if "A" =~ /\N{U+0041}/'
tab
snowman
A

\xNN takes exactly two hexadecimal digits. Braced \x{...} takes a hexadecimal code point of arbitrary length. For octal, prefer \o{...}; the older three-digit form can be confused with a numbered backreference. A braced escape also makes a pattern assembled from smaller strings easier to review.

Checkpoint: use a named or braced escape when a reader needs to know what character the pattern means. Keep the original input unchanged while testing, and print a clear success marker rather than relying on invisible whitespace.

4. Choose character classes with Unicode in mind

\d matches a decimal digit, \s whitespace and \w a word character. Their uppercase forms match the complement: \D, \S and \W. Perl can apply Unicode or locale rules, so these are not automatically ASCII-only. If the input format is explicitly ASCII, use the /a regular-expression modifier and test that decision.

\h and \v separate horizontal and vertical whitespace. \R matches a generic newline sequence, including a CRLF pair, and cannot be put inside a bracketed character class. \X matches an extended Unicode grapheme cluster, which is closer to a displayed character than a single code point.

$ perl -e 'print "ascii\n" if "A7" =~ /\A\w+\z/a; print "newline\n" if "\r\n" =~ /\A\R\z/; print "cluster\n" if "P\x{307}" =~ /\A\X\z/'
ascii
newline
cluster

The /a modifier is a deliberate boundary, not a performance switch. Without it, a pattern intended for identifiers may accept characters that your surrounding protocol does not define.

5. Quote literal input instead of hand-escaping it

When a variable contains text supplied by a user or a file, interpolate it through \Q and \E so punctuation is treated literally. This prevents characters such as parentheses, plus signs and brackets from becoming pattern operators.

$ perl -e 'my $literal = "(a+b)"; print "literal\n" if $literal =~ /\Q$literal\E/'
literal

\Q quotes until \E or the end of the pattern. It is not a substitute for validating the surrounding pattern, and it does not make an unsafe shell command safe. Keep the data in a Perl variable and do not build a shell command from it.

6. Make backreferences unambiguous

A capture stores text matched by parentheses. A backreference matches that same text again. Prefer the braced forms \g{1} for an absolute numbered reference, \g{-1} for a relative reference and \g{name} for a named reference. The named alternatives \k<name> and \k{name} are also supported.

$ perl -e 'print "duplicate\n" if "cat cat" =~ /(\w+) \g{1}/'
duplicate
$ perl -e 'print "palindrome\n" if "ABBA" =~ /(?<left>.)(?<right>.)\g{right}\g{left}/'
palindrome

The old \1 form is valid, but digits after it can make a dynamically assembled pattern ambiguous. A string ending in \g1 followed by another fragment beginning 37 can become \g137. Braces make the reference boundary explicit.

Do not assume \1 always means a backreference. Perl's old numeric syntax can be interpreted as octal depending on the number of capture groups and the digits used. Use \g{...} for a reference and \o{...} for octal when there is any doubt.

7. Anchor the input you actually mean

\A anchors the beginning of the string and \z anchors its true end. \Z also permits one trailing newline. These anchors are not changed by the /m modifier. Use \A...\z for a whole-string check when a final newline is not part of the format.

$ perl -e 'print "strict\n" if "id=42\n" =~ /\Aid=\d+\z/; print "trailing-newline-accepted\n" if "id=42\n" =~ /\Aid=\d+\Z/;'
trailing-newline-accepted

\b is a word/non-word boundary, not a line-start assertion. Its meaning depends on what \w considers a word character. For natural-language boundaries, Perl also provides Unicode boundary forms such as \b{wb} in versions that support them. Test a representative sample rather than assuming English ASCII behaviour.

8. Diagnose patterns without changing state

Start with a small literal string and print only whether the match succeeded. Add one escape at a time. If a warning says an escape is not recognised, read it as a compatibility warning rather than silently treating the spelling as stable.

$ perl -Mwarnings -e 'my $pattern = qr/\Q$ARGV[0]\E/; print "compiled\n"' 'literal (text)'
compiled

A failed match is not proof that the input is wrong. Check shell quoting, invisible newlines, the selected Unicode or ASCII mode, and whether a capture participated before its backreference. For a pattern assembled from fragments, print the final pattern with printf '%s\n' before compiling it. Never print secrets into a shared terminal or log just to debug a regex.

Done means

  • You checked the Perl version that will run the pattern.
  • You can distinguish a literal punctuation escape from a semantic escape.
  • You chose braced Unicode, octal and backreference forms where ambiguity matters.
  • You tested character classes with the intended Unicode or ASCII boundary.
  • You used \Q...\E for literal variable data and kept shell quoting separate.
  • You used \A and \z when a true whole-string match was required.