Choose Perl Character Classes Without Unicode Surprises
You will finish with Perl regular expressions that make their character set explicit: ASCII digits when an input format requires them, Unicode-aware classes when text is international, and bracketed sets when the allowed characters are narrow. The examples use Perl 5.38.2 and the matching perl-doc manual installed on this machine. Allow about fifteen minutes. You need Perl and a shell; no elevated privileges are required.
The route
Jump straight to the step you need, or tick off Done means at the end.
1. Check the Perl version and the manual
Character classes are regex components that consume one character. A dot, a backslash sequence such as \d, and a square-bracket class such as [aeiou] all belong to this family. Start by checking the interpreter that will run your code:
$ perl -v | sed -n '1,4p'
This is perl 5, version 38, subversion 2 (v5.38.2)
$ dpkg-query -W -f='${Package} ${Version}\n' perl-doc
perl-doc 5.38.2-3.2ubuntu0.6
$ man perlrecharclass
Checkpoint: if your program runs a different Perl than this shell, repeat the tests with that interpreter. The manual records version-specific changes, including extended bracketed classes becoming accepted in Perl 5.36 and vertical tab joining \s in Perl 5.18.
2. Decide whether a dot may cross a newline
By default, . matches one character other than a newline. The /s modifier changes that for the whole expression, while (?s:.) limits the change to one group. Test the distinction directly:
$ perl -e 'print "plain: ", ("\n" =~ /./ ? "match" : "no match"), "\n"; print "single-line: ", ("\n" =~ /./s ? "match" : "no match"), "\n"'
plain: no match
single-line: match
Do not use a dot as a vague substitute for an input policy. If a newline must never be accepted, keep the default or use \N, which the manual defines as a non-newline character unaffected by /s.
3. Choose ASCII or Unicode digits deliberately
\d matches one decimal digit. Without the ASCII modifier, it can match decimal digits from other writing systems. That is often correct for human text, but dangerous for identifiers, protocol fields and values that another system expects to contain only 0 to 9. The /a modifier restricts this class to ASCII:
$ perl -e 'my $arabic_indic = chr 0x0664; print "default: ", ($arabic_indic =~ /\d/u ? "match" : "no match"), "\n"; print "ASCII mode: ", ($arabic_indic =~ /\d/a ? "match" : "no match"), "\n"; print "ASCII 7: ", ("7" =~ /\d/a ? "match" : "no match"), "\n"'
default: match
ASCII mode: no match
ASCII 7: match
Use \d+ for a sequence of digits, not \d alone. If you accept Unicode digits, validate and interpret them as Unicode data rather than assuming that every glyph has the value its appearance suggests. For machine-facing decimal syntax, prefer /\d+/a and document that choice.
4. Match words and whitespace with the same care
\w means one word character, not a complete word. In ASCII mode it is the 63-character set [A-Za-z0-9_]. Without that restriction, Perl can include Unicode letters, marks and connector punctuation. Use \w+ only when that broader definition is intended.
Whitespace has several useful boundaries. \s is general whitespace, \h is horizontal whitespace and \v is vertical whitespace. Their complements are \S, \H and \V. The horizontal and vertical forms use Perl's native character set and do not follow the active locale, whereas \s, \d and \w can depend on the character-set rules in effect. This matters when a pattern is moved between a byte-oriented parser and Unicode text.
5. Build a small explicit class
Square brackets list alternatives and consume one character. A quantifier controls how many characters are consumed:
$ perl -e 'print "vowel: ", ("e" =~ /[aeiou]/ ? "match" : "no match"), "\n"; print "two vowels as one class: ", ("ae" =~ /^[aeiou]$/ ? "match" : "no match"), "\n"; print "two vowels with a quantifier: ", ("ae" =~ /^[aeiou]+$/ ? "match" : "no match"), "\n"'
vowel: match
two vowels as one class: no match
two vowels with a quantifier: match
Inside a normal class, most regex metacharacters lose their special meaning. Treat backslash, caret, hyphen, opening bracket and closing bracket as the exceptions: escape them when you mean the literal character. A leading caret negates the class, so [^0-9] means one character that is not an ASCII digit. It does not mean an empty string, and it does not validate a whole input unless you anchor and quantify the surrounding pattern.
6. Use POSIX classes when the name helps
POSIX classes sit inside a bracketed class. The delimiters are easy to misplace: [[:alpha:]] is correct, while [:alpha:] is just an ordinary class containing punctuation and letters. Useful names include alpha, digit, space, word and xdigit:
$ perl -e 'print "digit: ", ("7" =~ /[[:digit:]]/ ? "match" : "no match"), "\n"; print "not a digit: ", ("x" =~ /[[:^digit:]]/ ? "match" : "no match"), "\n"'
digit: match
not a digit: match
Perl extends POSIX syntax with a caret after the opening colon, as in [[:^digit:]]. Remember that classes combined inside one outer class are unioned. For example, a digit class combined with a non-hex-digit class matches almost everything except the ASCII letters a to f and A to F; it is not an intersection.
7. Intersect Unicode properties for a precise set
Unicode properties such as \p{Thai} and \p{Digit} describe sets independently. Perl's extended bracketed class, (?[ ... ]), can intersect, unite or subtract them. This feature was experimental from Perl 5.18 and accepted in Perl 5.36:
$ perl -e 'my $thai_digit = chr 0x0e57; print "Thai digit: ", ($thai_digit =~ /(?[ \p{Thai} & \p{Digit} ])/ ? "match" : "no match"), "\n"; print "ASCII 7: ", ("7" =~ /(?[ \p{Thai} & \p{Digit} ])/ ? "match" : "no match"), "\n"'
Thai digit: match
ASCII 7: no match
The operators are & for intersection, + or | for union, - for subtraction and ^ for symmetric difference. Whitespace is ignored inside this construct, but not between the opening (?[ or closing ]). Every character is treated as syntax, so write a literal a as [a]. Parenthesise mixed operations instead of relying on precedence.
These classes are compiled, not assembled at match time. Interpolate an already compiled extended class, such as qr/(?[ \p{Thai} + \p{Lao} ])/, rather than interpolating an unparsed string and assuming the visual grouping is preserved. A malformed class is a compile-time error, so run the exact pattern in a test before putting it into a long-running service.
Done means
- You checked the Perl interpreter and
perl-docversion used by the program. - You selected dot, backslash, bracketed, POSIX or extended syntax for a stated character set.
- You used
/awhen an input field must contain ASCII digits or word characters. - You added quantifiers and anchors when matching a complete number, word or field.
- You tested Unicode and extended classes with the same Perl version that will run in production.