Home / Alt manpages / perluniprops(1)

  • perluniprops(1)
  • User command
  • linux

Use Perl Unicode Properties Without Guessing the Syntax

You will finish with a small set of Perl commands that can recognise Unicode character properties in regular expressions and inspect a single code point with Unicode::UCD. The examples use Perl 5.38.2 on this machine, whose perluniprops(1) index describes Unicode Version 15.0.0. Allow about fifteen minutes. You need Perl and a shell; no root access or package changes are required.

This guide is about using the index, not memorising its long tables. The useful pattern is to choose a documented property, test it with a short expression, then use Unicode::UCD when you need an explanation of one character rather than a match against a string.

1. Check the Perl and Unicode versions

Start by recording the interpreter that will run your program:

$ perl -v
This is perl 5, version 38, subversion 2 (v5.38.2) built for x86_64-linux-gnu-thread-multi
$ perl -MUnicode::UCD -E 'say Unicode::UCD::UnicodeVersion()'
15.0.0

Your output may differ on another distribution. Keep the Unicode version with test results when matching behaviour across machines, because property tables and the set of matching code points can change with the Unicode data bundled into Perl.

Checkpoint

If the second command fails, the Perl installation does not provide the standard Unicode::UCD module needed later. The regular expression examples still use core Perl, but stop here before copying the inspection commands.

2. Match a property in a regular expression

Use \p{...} for characters that have a property and \P{...} for characters that do not. A compound property names the property and value, separated by = or :. This example accepts a string made entirely from Greek script characters:

$ perl -CS -Mutf8 -E 'for my $text ("Καλημέρα", "hello") { say "$text: ", ($text =~ /\A\p{Script_Extensions=Greek}+\z/ ? "match" : "no match") }'
Καλημέρα: match
hello: no match

Script_Extensions is useful when a character can be used with more than one script. The shorter \p{Greek} form is a Perl shortcut for this property in the documented cases, but the compound spelling tells the reader exactly which property is being tested.

For general text validation, \p{Word} is a Perl-defined property. It is not simply a synonym for one Unicode property:

$ perl -CS -Mutf8 -E 'for my $text ("cafe", "café", "東京", "123", "a space") { say "$text: ", ($text =~ /\A\p{Word}+\z/ ? "word" : "not one word") }'
cafe: word
café: word
東京: word
123: word
a space: not one word

That last result is a useful boundary: the expression checks the complete string because \A and \z anchor it. Without anchors, a search could succeed because only part of a longer string has the property.

3. Prefer stable compound names

The index lists Perl extensions and aliases as well as official Unicode properties. Some block shortcuts are explicitly discouraged because a future Unicode release could create a name collision or change what an alias means. Prefer the unambiguous compound form:

$ perl -CS -Mutf8 -E 'for my $text ("𐐀", "A") { say "$text: ", ($text =~ /\A\p{Block=Deseret}\z/ ? "Deseret block" : "other") }'
𐐀: Deseret block
A: other

Do not replace Block=Deseret with a casually chosen In_..., Is_... or bare block shortcut in a long-lived program. The manual identifies those forms as discouraged. The named compound property is clearer in code review and is the safer contract for a script that may run after a Perl or Unicode upgrade.

4. Inspect one code point with Unicode::UCD

Regular expressions answer whether text matches. Unicode::UCD::charprop answers what a Unicode property says about one code point. Pass the numeric code point and the documented property name:

$ perl -MUnicode::UCD -E 'say Unicode::UCD::charprop(0x41, "General_Category"); say Unicode::UCD::charprop(0x3B1, "Script_Extensions")'
Uppercase_Letter
Greek

0x41 is Latin capital A and 0x3B1 is Greek small alpha. The function is for Unicode properties, not every Perl-only extension. If you need all available property values for a code point, the same manual section points to charprops_all().

When you need the complete set of values for a property rather than the value for one character, use the appropriate Unicode::UCD function described by the index. prop_values("Script"), for example, returns the accepted values for that property on this installation:

$ perl -MUnicode::UCD -E 'say join ", ", (Unicode::UCD::prop_values("Script"))[0..4]'
Adlm, Aghb, Ahom, Arab, Armi

Do not assume that a property value is a character, a byte or a displayable glyph. Unicode properties describe code points and their classifications. Your input encoding and the way your program decodes text remain separate concerns.

5. Diagnose a surprising match

Check these points in order:

  • Confirm that the input is decoded text, not a byte string being treated as if it were Unicode.
  • Check the anchors. \A and \z test the whole string; an unanchored expression may only find a matching substring.
  • Check the property spelling and whether you used \p or \P. Changing the case of the letter before the brace reverses the test.
  • Check the installed Unicode version before comparing counts or edge cases with another machine.
  • Read the index entry for warnings marked deprecated, obsolete or discouraged. A deprecated property can produce a warning and may disappear in a future Perl release.

For a reproducible bug report, include the output of perl -v, the Unicode version, the exact property spelling and a short input that demonstrates the result. That gives someone else enough information to repeat the check without guessing at your environment.

Done means

  • You recorded the Perl and Unicode versions used by the check.
  • Your regex uses a documented \p{...} or \P{...} property and anchors where a whole-string test is intended.
  • You use stable compound names such as Script_Extensions=Greek or Block=Deseret where they make the contract explicit.
  • You can inspect a code point with Unicode::UCD::charprop instead of guessing what a match means.
  • You have not treated a Unicode property as an encoding conversion or a promise about how text will be displayed.