Trace Perl Regex Compilation with perlreguts
You will finish with a small Perl program that shows how a pattern becomes an internal regex program, how Perl looks for a possible start position, and how the matcher walks the resulting regops. This guide uses the installed Perl 5.38.2 and its perlreguts documentation. The output of the debugging facility is for investigation, not a stable interface for scripts.
The route
Jump straight to the step you need, or tick off Done means at the end.
Allow about twenty minutes. You need Perl, a shell and a pattern you understand. Everything here runs as your normal user and only prints diagnostics. No profile, service or system file is changed.
1. Check the interpreter you are studying
perlreguts(1) describes the implementation shipped with a particular Perl release. Start by recording the interpreter and the documentation version:
$ perl -v
This is perl 5, version 38, subversion 2 (v5.38.2) built for x86_64-linux-gnu-thread-multi
$ man perlreguts | col -b | sed -n '1,18p'
PERLREGUTS(1) Perl Programmers Reference Guide
Your architecture line and man-page header can differ. The useful checkpoint is that perl -v and the manual describe the same interpreter family. Do not treat a node name, structure field or function name in the manual as a promise that another Perl release will retain it.
2. Separate a pattern from its compiled program
The manual uses pattern for the source text you write and program for the internal representation Perl compiles from it. The program is a graph of regops rather than a simple tree. Literal text, branches, character classes and loop operations form nodes with links to later work.
This distinction explains why a short pattern does not map one-to-one to a short list of nodes. Perl can combine literal nodes, replace alternatives with a trie, remove construction-only nodes and attach optimisation metadata before matching. Read the listing as an implementation trace, not as a second regex syntax.
Checkpoint: verify that the ordinary match still has the result you expect before enabling diagnostics:
$ perl -e 'print "match\n" if "foobar" =~ /foo(?:bar|baz)/'
match
3. Dump compilation and one match
Load the re pragma in debug mode. It prints both compilation and execution information for regexes in that lexical scope:
$ perl -Mre=debug -e '"foobar" =~ /foo(?:bar|baz)/' 2>&1 | sed -n '1,32p'
Compiling REx "foo(?:bar|baz)"
Final program:
1: EXACT <fooba> (5)
5: TRIE-EXACT[rz] (12)
<r>
<z>
12: END (0)
anchored "fooba" at 0..0 (checking anchored) minlen 6
Matching REx "foo(?:bar|baz)" against "foobar"
Intuit: trying to determine minimum start position...
Intuit: Successfully guessed: match at offset 0
Match successful!
The exact trace is longer on this release and can change between releases. The significant observations are stable enough for a human investigation: the alternative was represented as TRIE-EXACT, the program has an END node, and the engine found an anchored literal before entering the main match loop.
Do not copy this diagnostic format into a test that must survive upgrades. The re documentation explicitly warns that debug output and its fine-grained modes are not an officially supported API.
4. Read the trace in two phases
perlreguts presents matching as two broad phases. Compilation parses the pattern into regops, links them, and analyses the result. Execution first tries to find a viable start position, then runs the interpreter from candidate positions.
In the example, EXACT <fooba> is a fixed string that gives the search a useful anchor. The minlen 6 line says a successful match cannot be shorter than six characters. These facts let Perl reject impossible input without running every operation at every position.
The execution section may show Intuit messages before numbered regops. That is the start-point analysis described by the manual, not a separate match result. The final Match successful! or failure message is the checkpoint for the whole operation.
5. Compare a pattern that cannot succeed
To see why start-point analysis matters, keep the pattern fixed and remove its required final character from the subject:
$ perl -Mre=debug -e '"abababababab" =~ /(a|b)*z/' 2>&1 | tail -n 12
Intuit: trying to determine minimum start position...
Intuit: Cannot find a start position
Match failed!
Freeing REx: "(a|b)*z"
The wording and the amount of output vary. The practical lesson is that a pattern with a required literal can provide a cheap rejection test, while a repeated alternation can still create substantial work on longer inputs. If you are investigating slow matches, capture a representative subject and pattern, then compare the trace with a simpler pattern rather than guessing from the source text alone.
6. Inspect more narrowly when the full trace is noisy
debug is deliberately broad. For a focused trace, use the capitalised Debug mode and select documented groups such as COMPILE, EXECUTE, PARSE, OPTIMISE or INTUIT:
$ perl -Mre=Debug,COMPILE -e 'my $re = qr/foo(?:bar|baz)/; print "compiled\n"' 2>&1 | sed -n '1,18p'
Compiling REx "foo(?:bar|baz)"
Final program:
1: EXACT <fooba> (5)
5: TRIE-EXACT[rz] (12)
<r>
<z>
12: END (0)
compiled
Keep the option spelling and grouping close to the re manual. The available detail and output are release-sensitive. If a mode is rejected, read perldoc re for the installed interpreter instead of importing an example from a different Perl version.
7. Keep implementation details out of production code
The manual names internal entry points such as re_intuit_start(), pregexec() and regmatch(), and structures such as regnode and regexp_internal. They are useful landmarks when reading Perl source, but they are not a supported application API. The manual also distinguishes positional REGNODE_AFTER links from execution-oriented regnext links, especially around branches and loops.
If you are writing an extension that replaces the regex engine, use perlreapi as the relevant interface document. If you are diagnosing an application, prefer normal Perl features and tests, and enable re diagnostics only in a local reproduction. Debug output can contain the pattern and subject, so redact sensitive data before sharing a trace.
Done means
- You recorded the Perl version before interpreting an internal trace.
- You can distinguish the source pattern from its compiled regop graph.
- You observed compilation, start-point analysis and execution separately.
- You used
rediagnostics only for human investigation. - You know that debug output, node names and structure details may change with Perl releases.
- You kept the examples unprivileged and did not alter persistent system state.