gendict compiles a UTF-8 word list into an ICU string trie dictionary, either a UCharsTrie or a BytesTrie, ready for a consuming application to load. The examples use the gendict supplied by Ubuntu's icu-devtools package, version 74.2-1ubuntu3.1 on the machine used here. Give it about ten minutes for a first build, including checking the input.
Install icu-devtools if gendict is not already on the box. Installing packages changes system state and normally needs elevated privileges, so do it through your normal package-management process rather than pasting in an unreviewed command. Confirm the executable and package version as an ordinary user:
command -v gendict
dpkg-query -W -f='${Package} ${Version}\n' icu-devtools
The source file must be UTF-8. Each useful line starts with a word and may then carry an integer value, decimal like 1 or hexadecimal with an ASCII 0x prefix like 0x2. The installed program reads the word from the start of the line and stops at the first whitespace. Keep the input deliberately simple while testing: malformed values and stray blank or whitespace-only lines produce diagnostics that are easy to miss.
Checkpoint: the version command is a useful first check, but this installed binary insists on a trie type in the same invocation. Run gendict --uchars -V or gendict --bytes -V rather than the bare gendict -V form.
Make a working directory and create an input file there. This example assigns one value per word; the values are data for whatever consumes the trie, gendict itself does not care what they mean:
mkdir -p "$HOME/tmp/gendict-demo"
cd "$HOME/tmp/gendict-demo"
cat > words.txt <<'EOF'
apple 1
banana 0x2
carrot 3
EOF
sed -n 'l' words.txt
A quoted here document keeps the example values literal. That final sed command makes tabs and trailing characters visible, which helps you catch an encoding or whitespace mistake before compilation. Validate a production word list's format in whatever tool generates it too. Do not add a line starting with spaces and assume it will be ignored: this installed 74.2 build reports an error for a line with no word, not a silent skip.
Choose exactly one output type. Use --uchars when the consumer expects a UChar string trie:
gendict --uchars --verbose words.txt words-uchars.dict
A successful run reports that it processed three lines, added three words, serialised a UCharsTrie, and wrote words-uchars.dict. The output is a binary ICU data file, not text, so check it exists and is non-empty:
test -s words-uchars.dict && file words-uchars.dict && stat -c '%s bytes' words-uchars.dict
Checkpoint: before running a command against an existing dictionary, inspect it and decide whether replacing it is really what you want. If a build goes wrong, remove or rename only the specific generated file, after checking the path. Writing to a new name, as above, keeps the whole exercise reversible and leaves any known-good file untouched.
A BytesTrie needs a transform. The common one shown by the manual is an offset transform: it subtracts a hexadecimal offset from input characters, so pick a value that maps every non-value character in your input into the byte range 0x00 through 0xFF. This example uses the offset from the installed help text:
gendict --bytes --transform offset-61 --verbose words.txt words-bytes.dict
On success, the verbose output identifies a BytesTrie and writes the second binary dictionary. Verify it independently:
test -s words-bytes.dict && stat -c '%s bytes' words-bytes.dict
cmp --silent words-uchars.dict words-bytes.dict; printf 'different output files: exit %s\n' "$?"
The two files should differ, because they hold different trie representations. Do not pick --bytes just to shrink the file: the reader must actually support that representation, and the transform has to match its decoding logic. The offset transform also gives special mappings to U+200D and U+200C, as the man page documents, so test those characters if your vocabulary uses them.
--uchars and --bytes is an error, and so is supplying both: the types are mutually exclusive.one, one that is too large, or extra non-whitespace text after it, can fail the build. Correct the source and rerun with a fresh output filename while you investigate, rather than forcing a bad file into use.If gendict cannot find required ICU data, set ICU_DATA to the directory containing it, trailing slash included where the install needs one, or pass --icudatadir DIRECTORY. On this install the normal data location sits under /usr/share/icu/74.2/; check the package contents before hard-coding a path on another machine:
dpkg -L icu-devtools | grep '/icu/74\.2/'
Only reach for --icudatadir when the error, or your deployment layout, actually calls for it: most shared-library installs never need it. If a failed or interrupted experiment already created an output file, do not deploy it. Replace it only after a clean run and a real read test from the application side.
gendict is the expected ICU version for the system it will ship on.--uchars or --bytes was used.