Home / Alt manpages / llvm-exegesis-20(1)

  • llvm-exegesis-20(1)
  • User command
  • linux

Measure LLVM Instruction Latency Safely with llvm-exegesis-20

You will finish with a repeatable way to generate an LLVM instruction snippet, run a controlled benchmark, save its YAML, and feed that result into the analysis mode. The examples use the installed llvm-exegesis-20 from Ubuntu package llvm-20, version 20.1.8, on this host's x86-64 Skylake target.

Allow 15 to 30 minutes for a first run. You need a shell, the LLVM 20 package, an instruction name supported by the selected target, and permission to use Linux performance counters. No command below needs root. Benchmarking executes generated or supplied machine code, so read the snippet and start with preparation-only or dummy-counter runs before trusting a real measurement.

1. Check the installed tool

First confirm that the command and package are the ones you intend to use. This is read-only:

$ command -v llvm-exegesis-20
/usr/bin/llvm-exegesis-20
$ llvm-exegesis-20 --version
Ubuntu LLVM version 20.1.8
  Default target: x86_64-pc-linux-gnu
  Host CPU: skylake
$ dpkg-query -W -f='${Package} ${Version}\n' llvm-20
llvm-20 1:20.1.8~++20250804090239+87f0227cb601-1~exp1~20250804210352.139

The host and package affect the result. A number measured on this machine is not a portable property of an instruction across CPUs. Check the local help when scripting, because the installed LLVM 20 binary exposes some options that are not present in older releases.

Checkpoint

If the version or target is not what you expect, stop here and fix the package selection before comparing results.

2. Generate a snippet without executing it

Use --benchmark-phase=prepare-snippet to inspect the generated instruction sequence while skipping measurement. This is a useful first pass and gives you YAML on standard output:

$ llvm-exegesis-20 \
    --mode=latency \
    --opcode-name=ADD64rr \
    --benchmark-phase=prepare-snippet
---
mode:            latency
key:
  instructions:
    - 'ADD64rr R14 R14 RDI'
  config:          ''
  register_initial_values:
    - 'R14=0x0'
cpu_name:        skylake
llvm_triple:     x86_64-pc-linux-gnu
min_instructions: 10000
measurements:    []
error:           actual measurements skipped.
info:            Repeating a single implicitly serial instruction
assembled_snippet: ''
...

The exact registers can change between runs or targets. The empty measurements list and the explicit skipped-measurements message are expected here. Do not treat this phase as a performance result. To generate every available opcode, the manpage documents --opcode-index=-1, but that can produce a large workload and is a poor first test.

Checkpoint

Confirm that the instruction, target and generated register setup are sensible before allowing code to execute.

3. Run a small benchmark and save YAML

For a smoke test that does not depend on usable hardware counters, add --use-dummy-perf-counters. This exercises snippet generation and execution but returns dummy counter values, so it is not evidence about latency:

$ llvm-exegesis-20 \
    --mode=latency \
    --opcode-name=ADD64rr \
    --use-dummy-perf-counters \
    --benchmarks-file=/tmp/add64rr-latency.yaml
$ printf 'exit status: %s\n' "$?"
exit status: 0
$ sed -n '1,24p' /tmp/add64rr-latency.yaml
---
mode:            latency
key:
  instructions:
    - 'ADD64rr RDX RDX R11'
  config:          ''
cpu_name:        skylake
llvm_triple:     x86_64-pc-linux-gnu
measurements:    []
error:           ''
...

--benchmarks-file writes benchmark YAML instead of sending it to standard output. A file named - means standard input or output, depending on the mode. The dummy option is useful for checking that a snippet does not crash, not for selecting a CPU model or estimating a cycle count.

For a real run, remove the dummy option and keep the output file:

$ llvm-exegesis-20 \
    --mode=latency \
    --opcode-name=ADD64rr \
    --benchmarks-file=/tmp/add64rr-latency.yaml

This may fail if LLVM was not built with libpfm support, the CPU is unsupported by libpfm, or the kernel restricts performance counters. A non-zero exit status and an error on standard error mean the result is not valid. Do not work around a counter permission failure by running an unfamiliar benchmark as root. Check the kernel's perf policy and the package build first. The command returns zero on success.

Checkpoint

Inspect the YAML and keep the host CPU, LLVM triple, mode and any error field with the measurement when you record results.

4. Compare uops and inverse throughput

The same opcode can be tested in the other measurement modes. Uops reports the generated micro-operation decomposition; inverse throughput estimates how quickly independent instances can be issued. They answer different questions from latency:

$ llvm-exegesis-20 --mode=uops --opcode-name=ADD64rr \
    --use-dummy-perf-counters --benchmarks-file=/tmp/add64rr-uops.yaml
$ llvm-exegesis-20 --mode=inverse_throughput --opcode-name=ADD64rr \
    --use-dummy-perf-counters --benchmarks-file=/tmp/add64rr-throughput.yaml

These commands are safe smoke tests for the same reason as the latency example, but dummy values must not be plotted or used to justify a scheduling change. For real comparisons, use the same binary, target settings, repetition settings and machine state for every run. The default target is the host; --mcpu can select a CPU model for work such as developing a scheduling model, but it does not turn the current host into that CPU.

5. Benchmark a custom assembly snippet

Use --snippets-file=- to read one from standard input. Start with an instruction that has no memory dependency:

$ printf '%s\n' 'vzeroupper' | \
    llvm-exegesis-20 --mode=uops --snippets-file=- \
    --benchmark-phase=prepare-snippet
---
mode:            uops
key:
  instructions:
    - 'VZEROUPPER'
  config:          ''
measurements:    []
error:           actual measurements skipped.
...

Custom snippets must have a valid register data flow. If a register is read before the snippet defines it, add an annotation such as LLVM-EXEGESIS-DEFREG XMM1 42. If a value should arrive from the benchmark setup, use LLVM-EXEGESIS-LIVEIN. On x86 Linux, the scratch-memory pointer is passed in RDI, so a memory snippet commonly begins with LLVM-EXEGESIS-LIVEIN RDI.

Do not paste an arbitrary assembly fragment into a real run. A bad instruction, invalid address or wrong annotation can terminate the benchmark process, and in-process mode is the default. Use --execution-mode=subprocess when you need memory annotations or an extra process boundary. That mode is restricted to x86-64 Linux in this tool.

6. Add memory only with explicit mappings

A snippet that dereferences a fixed address needs memory prepared at that address. The subprocess mode supports named definitions and mappings:

$ cat > /tmp/memory-snippet.s <<'EOF'
# LLVM-EXEGESIS-MEM-DEF test1 4096 7fffffff
# LLVM-EXEGESIS-MEM-MAP test1 8192

movq $8192, %rax
movq (%rax), %rdi
EOF
$ llvm-exegesis-20 --mode=uops --snippets-file=/tmp/memory-snippet.s \
    --execution-mode=subprocess --benchmark-phase=prepare-snippet

The definition creates a 4096-byte value and the mapping places it at decimal address 8192, which is hexadecimal 0x2000. The annotations are not a general memory safety mechanism. Review the address, size and instruction sequence together before measuring. If this example changes only a temporary file, undo it with:

$ rm -- /tmp/memory-snippet.s

The removal is destructive for that file, but it does not touch system configuration. The benchmark itself can still crash if the snippet reads outside the mapped range or writes unexpectedly.

7. Analyse saved benchmark results

Analysis consumes YAML from a latency or uops run and can write a CSV of performance clusters and an HTML report of inconsistencies with LLVM's scheduling information:

$ llvm-exegesis-20 --mode=analysis \
    --benchmarks-file=/tmp/add64rr-latency.yaml \
    --analysis-clusters-output-file=/tmp/add64rr-clusters.csv \
    --analysis-inconsistencies-output-file=/tmp/add64rr-inconsistencies.html
$ test -s /tmp/add64rr-clusters.csv && echo 'clusters written'
$ test -s /tmp/add64rr-inconsistencies.html && echo 'report written'

Analysis needs at least one of the two output options. Its default clustering algorithm is DBSCAN; use --analysis-clustering=naive when one cluster per opcode is more useful for investigating unstable results. Filter with --analysis-filter=reg-only or --analysis-filter=mem-only when memory behaviour would otherwise obscure the comparison.

Keep these reports beside the input YAML rather than treating them as independent facts. The selected triple and CPU are part of the benchmark context. If you deliberately analyse for a different combination with --mtriple or --mcpu, also pass --analysis-override-benchmark-triple-and-cpu and record that choice.

Done means

  • llvm-exegesis-20 --version identified the expected LLVM 20.1.8 binary and host target.
  • A preparation-only run was inspected before executing a generated snippet.
  • A real benchmark, or an explicitly labelled dummy-counter smoke test, produced YAML with a zero exit status.
  • Custom register and memory dependencies were annotated and reviewed before execution.
  • Analysis output, if used, is tied to the exact input YAML, CPU and LLVM triple.