Compile StringPrep Profiles into ICU Data with gensprep

Compile a filtered RFC 3454 StringPrep profile into ICU binary data with gensprep, without touching your source tree. You keep the destination explicit, so a bad run is easy to bin. Allow about 15 minutes if the ICU data tree and profile are already to hand.

This guide covers ICU 74.2, the version installed with the icu-devtools package on this machine.

1. Check the installed tool

Start without elevated privileges. This is a build tool, not a service, so there is no reason to use sudo when you can write to a working directory.

$ command -v gensprep
/usr/sbin/gensprep
$ dpkg-query -W -f='${Package} ${Version}\n' icu-devtools
icu-devtools 74.2-1ubuntu3.1
$ gensprep --help
Usage: gensprep [-options] [file_name]

Read the files specified and
create a binary file [package-name]_[bundle-name].spp with the StringPrep profile data

The installed executable prints a wider option set than the local gensprep(8) page. The examples below stick to the documented interface, the one stable contract in that page:

Do not assume an option works here just because another ICU release documents it.

2. Gather the filtered profile inputs

gensprep does not filter RFC 3454 data itself. It reads filtered files from the source tree. The man page says it looks under source/misc for files named rfc3454_*.txt, and under source/unidata for NormalizationCorrections.txt. Together they cover the unassigned set, mappings, prohibited code points and normalisation corrections a profile uses.

Point source at the root of an ICU data tree that contains those directories. Replace the placeholder with a real path.

Warning: Do not create empty files with the expected names to satisfy the check. That gives you incomplete data instead of a useful profile.

$ source_dir=/path/to/icu-data
$ test -d "$source_dir/misc" && test -d "$source_dir/unidata"
$ find "$source_dir/misc" -maxdepth 1 -type f -name 'rfc3454_*.txt' -print
$ test -f "$source_dir/unidata/NormalizationCorrections.txt"

The two test commands are checks only. A failed test stops the chain and tells you the tree is not ready. Find the correct ICU source or data checkout through your normal build process rather than substituting a system directory at random.

Checkpoint: The source directory contains misc and unidata, the RFC 3454 text files are present under misc, NormalizationCorrections.txt is present when the profile needs it, and the account that will run gensprep can read all of them.

3. Prepare a separate destination

The destination is where the generated binary data lands. Use a new, private build directory first. That avoids overwriting an existing ICU data file and makes a failed run trivial to remove.

$ build_dir=$(mktemp -d /tmp/gensprep-build.XXXXXX)
$ printf 'Using destination: %s\n' "$build_dir"
Using destination: /tmp/gensprep-build.XXXXXX

The printed suffix will differ on your system. Note it, or keep the shell open, because the next command uses the same variable. If you made the destination by hand, check that it is empty before you carry on:

$ find "$build_dir" -mindepth 1 -maxdepth 1 -print

If that prints unexpected existing output, stop and choose another directory.

Recovery: Do not use rm to clear a directory until you have checked its exact path and contents. The recovery action for this guide is to remove only the temporary directory after inspection: rm -rf -- "$build_dir".

4. Compile one profile

The executable takes a profile file name after its options. Use the filtered profile your build supplies, not a made-up name. The documented -s and -d options override the default location taken from ICU_DATA.

$ profile=/path/to/filtered-profile.txt
$ test -r "$profile"
$ gensprep --verbose --sourcedir "$source_dir" --destdir "$build_dir" "$profile"

A successful run returns to the shell without an error status and writes a generated binary profile into the destination. The installed help gives the output naming pattern as [package-name]_[bundle-name].spp. The exact name comes from the profile and package metadata, so do not hard-code a filename before the run.

If the command reports an unreadable file, a missing data directory or malformed profile data:

5. Inspect and hand off the result

List the destination and check that at least one output exists. The file is binary ICU data, so use metadata tools rather than an editor.

$ find "$build_dir" -maxdepth 1 -type f -printf '%f %s bytes\n'
$ find "$build_dir" -maxdepth 1 -type f -size +0c -print

The second command should print each non-empty generated file. Non-empty is a basic sanity check, not proof the profile is semantically correct. Have the consuming ICU build or test suite load the result before you copy it into a shared data archive. The man page records that ICU can read the binary directly, or that you can pass it to pkgdata(8) for inclusion in a larger archive.

When you are satisfied, copy the verified file into the build system's intended staging area, using its normal ownership and review process. That step may need elevated privileges in a system-owned directory. Compilation and inspection do not.

Tip: Keep the temporary output until the consumer has passed its checks.

Common traps

This guide changes only the temporary build directory.

Done means