Home / Alt manpages / llvm-bcanalyzer-18(1)

  • llvm-bcanalyzer-18(1)
  • User command
  • linux

Inspect LLVM 18 Bitcode Safely with llvm-bcanalyzer

You will use llvm-bcanalyzer-18 to inspect an LLVM bitcode file without modifying it. The normal report summarises the file's blocks, records, sizes and encoding; --dump adds a low-level, human-readable trace. Allow about ten minutes if you already have a bitcode file. The commands below are read-only apart from creating a temporary test file.

1. Check the installed tool

This guide targets Ubuntu's llvm-18 package, version 1:18.1.3-1ubuntu1. The executable reports LLVM 18.1.3. Check your own host before comparing output because block layouts and diagnostic details can vary between LLVM releases:

$ command -v llvm-bcanalyzer-18
/usr/bin/llvm-bcanalyzer-18
$ dpkg-query -W -f='${Package} ${Version}\n' llvm-18
llvm-18 1:18.1.3-1ubuntu1
$ llvm-bcanalyzer-18 --version
Ubuntu LLVM version 18.1.3
  Optimized build.

The installed command's name includes -18. Do not silently substitute an unversioned llvm-bcanalyzer from another LLVM installation when you need reproducible results.

2. Analyse an existing bitcode file

Pass one bitcode path as the final argument. The command writes its report to standard output and leaves the input unchanged:

$ llvm-bcanalyzer-18 /path/to/module.bc > /tmp/module-analysis.txt
$ printf 'exit status: %s\n' "$?"
exit status: 0
$ sed -n '1,24p' /tmp/module-analysis.txt
Summary of /path/to/module.bc:
         Total size: 11232b/1404.00B/351W
        Stream type: LLVM IR
  # Toplevel Blocks: 3

Per-block Summary:

The exact numbers depend on the module. The summary gives the total stream size in bits, bytes and 32-bit words, identifies the stream type, then lists each block's instances, size, records and abbreviations. The per-record histogram is useful when you are comparing two bitcode files or investigating why one is larger.

A zero exit status means the file was read successfully. It does not mean that the module is suitable for your compiler pipeline or that two modules are semantically equivalent. This tool measures the bitcode encoding; use LLVM's verifier or your normal build checks for semantic validation.

3. Read the low-level encoding

Add --dump when the block summary is not enough. Redirecting the output is sensible because a real module can produce a large trace:

$ llvm-bcanalyzer-18 --dump /path/to/module.bc > /tmp/module-dump.txt
$ sed -n '1,18p' /tmp/module-dump.txt
<IDENTIFICATION_BLOCK_ID NumWords=5 BlockCodeSize=5>
  <STRING abbrevid=4 op0=76 op1=76 op2=86 .../> record string = 'LLVM18.1.3'
  <EPOCH abbrevid=5 op0=0/>
</IDENTIFICATION_BLOCK_ID>
<MODULE_BLOCK NumWords=... BlockCodeSize=3>

The dump exposes block names, record names, abbreviations and operand values. It is a bitcode trace, not LLVM assembly. Use llvm-dis-18 when you need readable LLVM IR instead. Never treat numeric operand values as source-level names without checking the relevant LLVM bitcode format documentation.

In the installed LLVM 18 build, --help also lists --dump-blockinfo, --non-symbolic and --show-binary-blobs. These alter specialised dump output. Start with plain --dump and add one option at a time so that a captured trace remains understandable.

4. Feed bitcode through a pipeline

Omit the filename, or pass -, to read standard input. This lets you inspect a stream produced by another command:

$ cat /path/to/module.bc | llvm-bcanalyzer-18 - | sed -n '1,8p'
Summary of -:
         Total size: 11232b/1404.00B/351W
        Stream type: LLVM IR
  # Toplevel Blocks: 3

Per-block Summary:

The name in the report becomes - because the analyzer no longer has a path to display. For scripts, prefer an explicit temporary output file when you need an audit trail. A pipeline's exit status can hide an upstream failure in some shells, so check the producer separately or enable the shell's pipeline-failure option when writing a diagnostic script.

5. Distinguish bitcode from LLVM assembly

The analyzer expects bitcode, not a textual .ll file. If you have LLVM assembly, assemble a copy first with llvm-as-18:

$ llvm-as-18 /path/to/module.ll -o /tmp/module.bc
$ llvm-bcanalyzer-18 /tmp/module.bc | sed -n '1,8p'
Summary of /tmp/module.bc:
         Total size: ...
        Stream type: LLVM IR

Keep the temporary output separate from the source. The assembler creates a new bitcode file; the analyzer only reads it. If you accidentally pass the text file directly, the installed tool exits non-zero with an error such as Bitcode stream should be a multiple of 4 bytes in length. That is an input-format failure, not a reason to run the command with sudo.

6. Handle failures without losing evidence

Do not overwrite a report you may need for comparison. Use a new destination or a shell redirection guarded by a temporary file:

$ set -o pipefail
$ llvm-bcanalyzer-18 /path/to/module.bc > /tmp/module-analysis.new
$ status=$?
$ if [ "$status" -eq 0 ]; then
>     mv -- /tmp/module-analysis.new /tmp/module-analysis.txt
> else
>     rm -f -- /tmp/module-analysis.new
>     printf 'analysis failed with status %s\n' "$status" >&2
> fi

The mv runs only after a successful analysis, so an earlier report is retained on failure. The temporary report contains diagnostic data and can be removed after review. Do not use a broad wildcard with rm in a shared temporary directory.

Common distractions are simple: a zero-byte or non-bitcode input, a version mismatch between the producer and analyzer, and confusing the summary with a semantic verifier. Check the path, package version and exit status first. If the file came from an untrusted source, analyse it in a suitable isolated environment and avoid giving the command unnecessary filesystem access. The analyzer is intended to inspect data, but malformed input should still be treated as untrusted input.

Done means

  • You confirmed which LLVM 18 executable and package version you are using.
  • The analyzer read the intended bitcode file and returned status 0.
  • You can explain the total size, top-level blocks and per-block histogram in the report.
  • You used --dump only when a low-level trace was needed.
  • Standard input and textual LLVM assembly are handled deliberately.
  • Existing reports and source files remain intact when analysis fails.