Home / Alt manpages / perlfaq4(1)

  • perlfaq4(1)
  • User command
  • linux

Sort Perl Hashes by Key or Value Without Surprises

You will turn a Perl hash into a predictable report, sorted either by its keys or by its values. The examples use only core Perl and take about ten minutes to try. They were checked with Perl 5.38.2, the version installed with this guide's perl-doc package.

Checkpoint

Start at step 1 if you are unsure whether your data is text or numeric. Jump to step 4 if the report already exists and only its ordering is wrong.

1. Make the data and inspect the default trap

A hash has no useful insertion order for reporting. Start with a small, self-contained data set so the ordering rule is visible:

use strict;
use warnings;

my %score = (
    alice => 17,
    bob   => 4,
    carol => 12,
);

for my $key (keys %score) {
    printf "%s %d\n", $key, $score{$key};
}

Save that as /tmp/hash-report.pl and run perl /tmp/hash-report.pl. The three lines may appear in an order you did not expect. That is normal: keys %score supplies the keys, but does not sort them.

Do not build an operational report around the apparent order from a single run. A hash's internal arrangement is not a presentation rule.

2. Sort by key when names are the report order

Pass the keys through sort before looking up their values:

for my $key (sort keys %score) {
    printf "%s %d\n", $key, $score{$key};
}

The expected output is:

alice 17
bob 4
carol 12

Plain sort uses string comparison. That is appropriate for names and labels, and is also why sorting numeric-looking keys can surprise you: text order puts 10 before 2. If the keys represent numbers, supply a numeric comparator instead:

my %quantity = (2 => 8, 10 => 3, 1 => 9);

for my $key (sort { $a <=> $b } keys %quantity) {
    printf "%d %d\n", $key, $quantity{$key};
}

Here <=> compares numbers; cmp compares strings. Choose deliberately rather than converting the values only when printing.

3. Sort by value and make ties repeatable

To rank the entries, still sort a list of keys, but compare the values reached through those keys:

my @by_score = sort { $score{$a} <=> $score{$b} } keys %score;

for my $key (@by_score) {
    printf "%s %d\n", $key, $score{$key};
}

This prints the lowest score first:

bob 4
carol 12
alice 17

Equal values need a second comparison if stable, reviewable output matters. The or expression uses the key only when the numeric comparison ties:

my @by_score_then_name = sort {
    $score{$a} <=> $score{$b}
        or $a cmp $b
} keys %score;

For descending order, reverse the numeric operands: $score{$b} <=> $score{$a}. Keep the key comparison as the tie-breaker unless you have a documented reason to accept arbitrary order.

4. Avoid expensive work inside the comparator

The block passed to sort can run many times for the same element. If comparing requires a costly calculation, compute the sort key once and retain the original record. For simple values, the direct examples above are clearer. For case-insensitive labels, this is enough:

my @names = sort { lc($a) cmp lc($b) } keys %score;
print "$_\n" for @names;

For a large hash or an expensive expression, use a decorated list so the calculation is cached:

my @ordered = map  { $_->[0] }
              sort { $a->[1] cmp $b->[1] }
              map  { [$_, lc($_)] } keys %score;

This is a memory-for-CPU trade-off. The intermediate list contains every key and its cached value, so measure before using it on a very large hash.

5. Choose keys or each for the size and order you need

sort keys %score creates a list of keys. That is the right choice when ordering matters, and usually the simplest choice for a report. If the hash is very large and order does not matter, each visits one key-value pair at a time:

while (my ($key, $value) = each %score) {
    printf "%s %d\n", $key, $value;
}

The output order from each is not a report order. Do not mix each with keys, values or another each loop on the same hash while relying on the iterator: those operations can reset it. Also avoid adding or deleting unrelated keys during an each loop, because Perl may rearrange the hash and skip or revisit entries.

If you need to change the hash, collect the keys first, then modify it in a separate loop. That makes the set being processed explicit and avoids iterator surprises.

Done means

  • The report uses sort keys %hash when key order is required.
  • Numeric data uses <=>; text uses cmp.
  • Value-sorted output has a deliberate tie-breaker where repeatability matters.
  • You use each only when arbitrary order and lower temporary memory are acceptable.
  • Running perl /tmp/hash-report.pl produces the expected report without changing any system file.