Measure x86 Instruction Latency with llvm-exegesis-18
You will measure an LLVM instruction on the current CPU, save the YAML result, and inspect the generated machine-code snippet. This guide uses Ubuntu's llvm-exegesis-18 package, version 18.1.3, on a 64-bit x86 Linux host. Allow about 10 minutes for a first useful comparison. You need a shell and the package already installed; the examples do not need root.
The route
Jump straight to the step you need, or tick off Done means at the end.
What the tool measures
llvm-exegesis-18 takes an LLVM opcode and a mode, generates a small sequence intended to make the relevant behaviour measurable, then reports YAML. Latency makes execution as serial as possible. uops measures the instruction's micro-operation decomposition, while inverse_throughput measures how quickly independent work can be issued. The result describes the machine on which you ran it, not a universal property of the instruction.
The installed binary reports LLVM 18.1.3 and identifies this host as x86_64-pc-linux-gnu with a Skylake host CPU. Benchmarking modes are supported on x86-64, AArch64, MIPS and PowerPC64LE Linux, but some functions are architecture-specific. Analysis mode is broader because it reads existing results rather than executing a new snippet.
1. Confirm the binary and host
Start by checking the exact binary selected by your shell. This prevents a result from silently coming from another LLVM release.
command -v llvm-exegesis-18
llvm-exegesis-18 --version
Expected output begins like this, although the target and CPU depend on the machine:
/usr/bin/llvm-exegesis-18
Ubuntu LLVM version 18.1.3
Checkpoint
Continue only if the command exists and its reported target matches the machine you intend to study. The tool assembles and may execute code for the selected subtarget, so do not treat it as a cross-platform simulator.
2. Run a small latency measurement
Measure the latency of the x86 ADD64rr opcode and store the YAML in a temporary file. The dummy counter option is useful for a first smoke test or a system where hardware performance counters are unavailable. It returns a dummy measurement, so remove that option when you need a real counter-based result.
llvm-exegesis-18 \
--mode=latency \
--opcode-name=ADD64rr \
--use-dummy-perf-counters \
--num-repetitions=1000 \
--benchmarks-file=/tmp/add64rr-latency.yaml
Check that the command succeeded and that the output identifies the opcode and mode:
test -s /tmp/add64rr-latency.yaml && \
sed -n '1,18p' /tmp/add64rr-latency.yaml
A successful file contains fields such as mode: latency, an instruction line containing ADD64rr, cpu_name, llvm_triple, num_repetitions, and measurements. The generated register assignment can change between runs. A zero exit status means the tool completed; it does not make a dummy value a hardware measurement.
For a real measurement, run the same command without --use-dummy-perf-counters. The operating system and perf security settings may prevent access to counters. Treat an error as a system capability problem, not as evidence about the instruction. Do not work around such a restriction by changing security policy just for a quick benchmark.
3. Inspect generated code without executing it
Before testing an unfamiliar opcode or snippet, use a preparation phase. It generates the sequence but stops before measurement. This is a useful review point for register choices and for catching an unsupported opcode without running the generated code.
llvm-exegesis-18 \
--mode=latency \
--opcode-name=ADD64rr \
--benchmark-phase=prepare-snippet \
> /tmp/add64rr-prepared.yaml
sed -n '1,18p' /tmp/add64rr-prepared.yaml
Look for actual measurements skipped. in info. The assembled_snippet field is empty at this phase. The later phases, prepare-and-assemble-snippet and assemble-measured-code, expose progressively more of the generated sequence; measure runs it.
4. Compare throughput modes
Use the same opcode and output convention when comparing modes. These commands write separate YAML files so one result cannot overwrite another.
llvm-exegesis-18 --mode=uops \
--opcode-name=ADD64rr \
--use-dummy-perf-counters \
--benchmarks-file=/tmp/add64rr-uops.yaml
llvm-exegesis-18 --mode=inverse_throughput \
--opcode-name=ADD64rr \
--use-dummy-perf-counters \
--benchmarks-file=/tmp/add64rr-throughput.yaml
--num-repetitions is a target number of executed instructions, not necessarily the number of times the whole snippet is copied. The actual repetition count is divided by snippet size. Higher values can reduce noise but take longer. The default repetition mode is duplicate; loop and min change how the generated work is arranged. Keep these settings identical when the comparison is meant to isolate an opcode.
5. Benchmark a custom assembly snippet
For code that the opcode interface cannot express, provide assembly through standard input or a file. This harmless example measures the uops mode for vzeroupper:
printf '%s\n' 'vzeroupper' | \
llvm-exegesis-18 \
--mode=uops \
--snippets-file=- \
--use-dummy-perf-counters \
--benchmarks-file=/tmp/vzeroupper-uops.yaml
The tool checks that register uses have a definition or a live-in annotation. A scratch-memory pointer is available through the platform's designated register when that register is marked live in. On x86 Linux the manpage gives RDI as the example. A value can be supplied with LLVM-EXEGESIS-DEFREG:
# LLVM-EXEGESIS-LIVEIN RDI
# LLVM-EXEGESIS-DEFREG XMM1 42
vmulps (%rdi), %xmm1, %xmm2
vhaddps %xmm2, %xmm2, %xmm3
addq $0x10, %rdi
Memory annotations require subprocess execution. LLVM-EXEGESIS-MEM-DEF creates a named byte pattern and LLVM-EXEGESIS-MEM-MAP places it at a decimal address. This is an execution setup, not ordinary process memory you should assume is safe to reuse. Review every address and instruction before running it.
6. Analyse saved results
Analysis mode reads a benchmark YAML file and can write clusters as CSV or scheduling inconsistencies as HTML. It does not need a new measurement. At least one analysis output option is required.
llvm-exegesis-18 \
--mode=analysis \
--benchmarks-file=/tmp/add64rr-latency.yaml \
--analysis-clusters-output-file=/tmp/add64rr-clusters.csv \
--analysis-inconsistencies-output-file=/tmp/add64rr-inconsistencies.html
Verify both files before opening or processing them:
test -s /tmp/add64rr-clusters.csv
test -s /tmp/add64rr-inconsistencies.html
head -n 4 /tmp/add64rr-clusters.csv
Clusters group results with similar performance characteristics. The inconsistency report compares those observations with LLVM scheduling information. If the binary was not built with debug information, scheduling class names may be shown as numeric IDs; that does not invalidate the analysis.
Common traps and clean-up
- Do not compare a result labelled for one CPU with a result from another CPU and call the difference an LLVM change. Record
cpu_name, target triple, mode, repetition settings and LLVM version. - Do not use
--opcode-index=-1casually. It measures every existing opcode and can take substantially longer than one instruction. - Benchmarking is not read-only in the CPU sense. The default execution mode is in-process and the generated snippet is normally executed. Use a preparation phase first, and use subprocess mode when memory annotations require it.
- The commands above create only temporary files. Remove these exact files when you no longer need them:
rm -f /tmp/add64rr-latency.yaml /tmp/add64rr-prepared.yaml /tmp/add64rr-uops.yaml /tmp/add64rr-throughput.yaml /tmp/vzeroupper-uops.yaml /tmp/add64rr-clusters.csv /tmp/add64rr-inconsistencies.html. This is optional and does not alter the installed package.
Done means
llvm-exegesis-18 --versionreports the intended LLVM release and target.- A YAML file records the selected opcode, mode, CPU and repetition count.
- You have distinguished a dummy-counter smoke test from a real hardware measurement.
- You inspected or prepared custom snippets before executing unfamiliar code.
- Analysis outputs exist and are tied to the same benchmark input.