Find Printable Text in Binaries with llvm-strings-18
You will finish with a repeatable way to pull readable ASCII sequences from a binary or other byte stream, adjust the minimum length, and add file names or byte offsets when the result needs investigation. The examples use llvm-strings-18 from Ubuntu package llvm-18, version 1:18.1.3-1ubuntu1 on the machine used for this guide.
The route
Jump straight to the step you need, or tick off Done means at the end.
Allow about ten minutes. You need a shell and a file you are allowed to inspect. No elevated privileges are normally needed. Reading a file can still expose secrets, tokens or personal data, so treat the output as sensitive and do not paste it into tickets or chat without checking it first.
1. Confirm the installed command
Check that the versioned executable is the one on your path. This is an ordinary read-only check:
$ command -v llvm-strings-18
/usr/bin/llvm-strings-18
$ llvm-strings-18 --version
llvm-strings-18
Ubuntu LLVM version 18.1.3
Optimized build.
The manpage describes this program as a drop-in replacement for GNU strings, but its default scan is worth remembering: it searches the entire input, regardless of file format. It is not restricted to recognised object-file sections.
Checkpoint: if command -v finds nothing, install or select the LLVM package that provides the command according to your distribution's normal package process. Do not substitute an unverified binary merely because it has the same name.
2. Scan a file with the default threshold
Pass one or more input paths after the options. A printable string is a run of printable ASCII characters. The default minimum is four characters, and any other byte, including the end of the file, ends the run:
$ llvm-strings-18 ./app.bin
GLIBC_2.34
usage: app [options]
build 2026-09-24
...
The exact lines depend on the file. Short fragments, such as two-letter identifiers, will not appear with the default threshold. This makes the default useful for a quick inspection, but it is not a complete text extractor and it does not decode arbitrary character encodings.
Multiple paths are accepted. The output is combined, which is convenient for a quick comparison but can make provenance unclear. Add file names when that matters in the next step.
3. Lower the minimum for short markers
Use -n or --bytes to set the minimum number of printable ASCII characters. The value is a length, not a byte offset:
$ llvm-strings-18 --bytes=3 ./app.bin
API
foo
CFG
$ llvm-strings-18 -n 6 ./app.bin
GLIBC_2.34
usage:
build 2026-09-24
A smaller threshold increases noise quickly, especially in compiled data. Start with the default, then lower it only when you have a reason to look for short markers. If you need a different encoding or structured metadata, use a tool designed for that format instead of assuming every printed sequence is meaningful text.
Checkpoint: verify the threshold against a controlled stream before interpreting a large file:
$ printf 'ab\000abcd\000' | llvm-strings-18 -n 3 -
abcd
The single hyphen explicitly means standard input. With no input path at all, llvm-strings-18 also reads standard input.
4. Keep results tied to their input files
Use -f or --print-file-name when scanning more than one file, or whenever you are recording evidence:
$ llvm-strings-18 --print-file-name ./app.bin ./libhelper.so
./app.bin: usage: app [options]
./libhelper.so: GLIBC_2.34
./libhelper.so: helper_init
The prefix is the path you supplied. It is not a claim about which object-file section contained the string. Because LLVM scans the whole input, the result can come from headers, sections, padding or embedded data.
The compatibility option -a or --all is silently ignored. It exists for GNU strings command lines; it does not change this program's whole-file behaviour.
5. Add offsets for binary investigation
Use -t or --radix to print the offset before each string. The permitted radix values are d for decimal, o for octal and x for hexadecimal:
$ llvm-strings-18 -t x ./app.bin
40 usage: app [options]
2a8 build 2026-09-24
$ llvm-strings-18 --radix=d ./app.bin
64 usage: app [options]
680 build 2026-09-24
Offsets are useful when you want to correlate a string with a hex dump or a later parser. They are offsets within the input file, not virtual addresses and not source-code line numbers. Keep the radix in your notes so another person can reproduce the lookup.
If the radix is invalid, the command exits non-zero. Treat that as a command error rather than silently reading an incomplete result:
$ llvm-strings-18 -t q ./app.bin
llvm-strings-18: error: invalid radix 'q'
$ printf 'exit status: %s\n' "$?"
exit status: 1
6. Use standard input in a pipeline
For a stream or a command that produces bytes, pipe into llvm-strings-18 and use - as the input marker when you want the boundary to be explicit:
$ dd if=./app.bin bs=1M status=none | llvm-strings-18 -n 8 -
usage: app [options]
build 2026-09-24
The command does not modify its input. There is no undo step for these examples because they only read bytes and write extracted text to standard output. Redirect output to a file only after checking the destination, particularly if it may overwrite an existing investigation record.
For an untrusted or very large file, consider the practical cost before running the scan. The tool reads the whole input, and a low threshold can produce a large output. Use shell redirection or a pager deliberately, and protect any output file with the same access controls as the source.
Done means
llvm-strings-18 --versionidentified the installed LLVM 18 executable.- You scanned a file with the four-character default and understood that the whole file is searched.
- You used
-nonly when short strings were relevant. - You used
-fto preserve provenance and-t d,-t oor-t xfor reproducible offsets. - You treated extracted strings as potentially sensitive and made no persistent system change.