Home / Alt manpages / perlperf(1)

  • perlperf(1)
  • User command
  • linux

Measure Perl Hot Spots with Benchmark and Safe Profiling

You will finish with a repeatable Perl performance check: measure a working program, identify where it spends time, make one controlled change, and verify that it still produces the same result. The local perlperf manual is a guide, not an executable command. These examples use Perl 5.38.2 and the perl-doc package installed on this machine.

Allow about 20 minutes for a small script, plus however long your real test data takes. You need Perl, a test or sample input, and a version-control checkout if you are changing application code. Do not optimise an untested program. A faster incorrect result is still a failure.

1. Confirm the tools and make a baseline

Check the interpreter and documentation package with ordinary, read-only commands:

$ perl -v
This is perl 5, version 38, subversion 2 (v5.38.2)
$ dpkg-query -W -f='${Package} ${Version}\n' perl-doc
perl-doc 5.38.2-3.2ubuntu0.6
$ command -v time
/usr/bin/time

Run the program once with a known input and record its output. Then measure the same invocation with the shell's time utility:

$ time perl bin/report.pl data/sample.txt > /tmp/report.before

real    0m0.005s
user    0m0.003s
sys     0m0.001s

Your figures will differ. real is elapsed wall-clock time, including waiting. user is CPU time spent by the program in user space, and sys is CPU time spent in the kernel on its behalf. Keep the input, command, Perl version and output together. Repeating the run matters because other processes can affect timings.

Checkpoint: the baseline is useful only if the program's output is correct and the command can be repeated. Save a checksum or a test result before editing anything:

$ cksum /tmp/report.before
3044173777 52029194 /tmp/report.before
$ perl bin/report.pl data/sample.txt > /tmp/report.repeat
$ cksum /tmp/report.repeat

Do not treat the sample checksum as a universal expected value. It is only an example of recording your own result.

2. Compare a small operation with Benchmark

The Benchmark module is suitable when you are comparing a focused piece of Perl code. Give each candidate the same inputs and return a value from the code under test, so the comparison measures real work:

$ perl -MBenchmark=timethese -e '
my $source = "0123456789abcdefghijklmnopqrstuvwxyz";
timethese(100000, {
    regex => sub { my $s = $source; $s =~ s/[aeiou]/x/g; return $s },
    tr    => sub { my $s = $source; $s =~ tr/aeiou/xxxxx/; return $s },
});'
Benchmark: timing 100000 iterations of regex, tr...

The report includes wall-clock and CPU measurements, an iteration count, and a calculated rate. The exact numbers depend on this machine and may be too small to display cleanly for a short test. Increase the iteration count until the result settles, but keep each candidate equivalent. Here tr is a fixed character translation, while the regular expression is more flexible. A faster operation is not automatically the right replacement if it changes the required behaviour.

Checkpoint: run the benchmark twice and compare the ordering, not just a single rate. If the order changes or the figures are noisy, use more iterations or a more representative input. Never infer an application-wide improvement from a microbenchmark alone.

3. Find the hot path with a profiler

Use a profiler after you have a repeatable baseline. The local manual shows the Perl debugger switch with Devel::NYTProf and reports from nytprofhtml or nytprofcsv. First check whether that module and its report tools are installed:

$ perl -MDevel::NYTProf -e 'print "$Devel::NYTProf::VERSION\n"'
Can't locate Devel/NYTProf.pm in @INC ...
$ command -v nytprofhtml || true

On this installation, NYTProf is not available, so do not copy the profiling command and expect it to work. Installing extra packages is an administrator or project-owner decision outside this guide. The same check avoids confusing a missing module with a bug in your program.

If your approved environment already provides NYTProf, profile a normal workload into a separate working directory:

$ mkdir -p /tmp/perlperf-run
$ cd /tmp/perlperf-run
$ perl -d:NYTProf /path/to/bin/report.pl /path/to/data/sample.txt > report.out
$ nytprofhtml

Use a disposable directory because the profiler writes a report database, normally nytprof.out, and the HTML command creates a report tree. The generated report is diagnostic data and can expose file names, query text or other sensitive details. Keep it private and remove it only after you no longer need it. Do not profile production traffic casually: instrumentation adds overhead and can increase storage use.

Read the report from the top down. A subroutine with high exclusive time is doing work itself; high inclusive time can include expensive callees. A line or subroutine called many times can be a better target than a rare, slow call. Confirm the observation with a second representative run before changing code.

4. Change one thing and verify the same task

Make one narrow edit in version control. The perlperf guidance calls out familiar trade-offs: a hash lookup can avoid repeatedly scanning a list, built-in cmp and <=> sorting is usually cheaper than a custom comparison, and caching sort keys can cost memory and an extra pass. These are hypotheses, not rules. Measure your data.

Before replacing the old code, preserve its output and run the test suite:

$ git diff --check
$ prove -lr t
All tests successful.
$ time perl bin/report.pl data/sample.txt > /tmp/report.after
$ cksum /tmp/report.before /tmp/report.after
3044173777 52029194 /tmp/report.before
3044173777 52029194 /tmp/report.after

If the checksums differ, stop. A changed output may be an intended correction, but it is not evidence of a safe performance improvement. Inspect the difference and restore the change with your normal version-control workflow if it was not intended. Do not use a destructive reset when you have uncommitted work you may need.

Run the same measured command several times after the edit. Compare like with like: same input, environment, output destination and warm-up conditions. Keep both the timing and correctness result. A lower real time with a much higher sys time may move work into the kernel rather than remove it.

5. Avoid the common measurement traps

  • Do not optimise before correctness and tests. Otherwise a benchmark rewards a broken shortcut.
  • Do not change several functions at once. You lose the link between the edit and the measured result.
  • Do not compare different inputs, output paths or debug settings. Logging and diagnostics can dominate a short run.
  • Do not trust a tiny benchmark result. The module warns when there are too few iterations for a reliable count.
  • Do not enable verbose debug logging while measuring normal work. Constructing a message can cost time even when the logger later discards it; guard expensive diagnostic work before calling the logger.
  • Do not run a profiler as root merely to make a report. Use the least privilege needed to read the input and write the report.

Done means

  • The program had a correct, repeatable baseline and a recorded input.
  • Any microbenchmark used equal work and enough iterations to produce stable results.
  • A profiler, where available, identified a measured hot path rather than a guess.
  • Only one change was made between comparisons.
  • The edited program passed its tests, produced the same required output, and improved the measured workload.
  • Temporary reports and sensitive inputs remain private, and no production service was changed during measurement.