Home / Alt manpages / perldbmfilter(1)

  • perldbmfilter(1)
  • User command
  • linux

Keep DBM Storage Rules Out of Your Perl Code with Filters

Perl DBM filters let you transform keys and values at the boundary of a tied hash. The application can use ordinary strings while the database receives the representation required by another program, such as a C application expecting terminating NULL bytes. This guide shows the four filter hooks, a safe working example, and the boundaries that matter when you read or change an existing database.

Prerequisites: Perl 5.36 or newer is used in the examples, with SDBM_File available. The installed reference on this system is Perl 5.38.2 from the perl-doc package. Allow about 15 minutes for a new filter and a little longer if you are adapting an existing database format. No elevated privileges are needed for the example because it works in a temporary directory.

1. Choose the hook that matches the boundary

All five DBM modules shipped with this Perl provide the same four methods: DB_File, GDBM_File, NDBM_File, ODBM_File and SDBM_File. A store filter runs when Perl writes a key or value. A fetch filter runs when Perl reads one.

  • filter_store_key changes a key before it is written.
  • filter_store_value changes a value before it is written.
  • filter_fetch_key changes a key after it is read.
  • filter_fetch_value changes a value after it is read.

Start with the smallest pair that expresses the format. If only keys need packing, install key filters and leave values alone. Each hook holds one filter at a time. Calling a method returns the filter that was installed previously, or undef if there was none.

Checkpoint: filters operate on $_

When Perl calls a filter, the local $_ contains the key or value being processed. Change $_ in place. The filter's return value is ignored, so a filter such as sub { $_ .= "\0" } is correct even though it does not return the modified string explicitly.

2. Add and remove terminating NULL bytes

This complete example stores a trailing NULL byte for both keys and values, then removes it again on fetch. That lets the main code use abc and def while a C consumer sees the terminated byte strings.

use v5.36;
use SDBM_File;
use Fcntl qw(O_CREAT O_RDWR);

my $filename = "/tmp/perl-filter-example";
my %hash;
my $db = tie(%hash, 'SDBM_File', $filename, O_RDWR|O_CREAT, 0600)
  or die "Cannot open $filename: $!\n";

$db->filter_fetch_key(sub { s/\0$// });
$db->filter_store_key(sub { $_ .= "\0" });
$db->filter_fetch_value(sub {
    no warnings 'uninitialized';
    s/\0$//;
});
$db->filter_store_value(sub { $_ .= "\0" });

$hash{abc} = 'def';
die "unexpected value\n" unless $hash{abc} eq 'def';
print "read=$hash{abc}\n";

undef $db;
untie %hash;

Expected output is:

read=def

The s/\0$// operation removes one NULL at the end. The no warnings 'uninitialized' scope is useful for a fetch filter because a missing or unusual record may not provide a defined value to transform. It does not repair malformed data, and it should not be used to hide errors in the main application.

The example creates database files under /tmp. Once you have finished testing, remove the exact files it created, for example with rm -f /tmp/perl-filter-example /tmp/perl-filter-example.dir /tmp/perl-filter-example.pag. Check the names first on your DBM implementation. That command is destructive for those files, so do not run it against a real database.

3. Verify the stored representation before sharing it

A filter changes the bytes crossing the DBM interface; it does not convert every database already on disk. Test the writer and reader together first. Then inspect the database with the other application or a purpose-built diagnostic that understands its format. A plain text dump can be misleading because a terminating NULL is not normally visible.

Keep the filters symmetrical. If a store filter appends a NULL, the matching fetch filter should remove it. Applying the store filter twice would append two bytes, while applying the fetch filter to an un-terminated value can remove a meaningful final character only if the format assumption is wrong. Make the format contract explicit in comments and tests.

Checkpoint: pack numeric keys only when the format requires it

DBM normally stores Perl hash keys and values as strings. If an external program requires a key to be a C integer, pack it before storage and unpack it after fetching:

use DB_File;

$db->filter_fetch_key(sub { $_ = unpack("i", $_) });
$db->filter_store_key(sub { $_ = pack("i", $_) });

$hash{123} = 'def';

This is a binary interface, not a general portability layer. The C integer size, byte order and signedness must agree between both programs. Test on every architecture you support. Do not add value filters unless the value has the same external representation.

4. Remove a filter without closing the database

Pass undef to the relevant method to uninstall its filter:

my $old_filter = $db->filter_store_value(undef);

The call returns the filter that was removed. If you need to restore it temporarily, keep that return value and pass it back later. Uninstalling a filter does not rewrite records already stored with the transformed representation. After changing the rules, read and write only data whose representation you have deliberately checked.

Common traps

  • Do not assume a fetch filter's return value is used. Modify $_.
  • Do not install the same transformation in both directions. Store and fetch operations need inverse transformations.
  • Do not mix filtered and unfiltered processes against one database unless both representations are documented and tested.
  • Do not use filters as encryption or access control. They provide representation changes, not confidentiality or authentication.
  • Do not overwrite an existing database while experimenting. Copy it first and confirm that the DBM implementation's companion files are included.

Done means

  • The chosen store and fetch hooks match the key or value boundary you need.
  • A write followed by a read returns the application-level value you expect.
  • The external consumer has verified the actual bytes, including NULL termination or integer layout.
  • Missing records and malformed data produce a deliberate result rather than a hidden warning.
  • You can name the exact database files to recover or remove if the experiment goes wrong.