Compilers rarely tell you why one instruction sequence is slower than another that looks the same, but llvm-mca-20 will. Feed it a short x86 assembly sequence and it turns that into a repeatable report of simulated cycles, instructions per cycle, instruction latency, and pressure on the processor's execution resources. This is a static scheduling-model analysis, not a benchmark of your actual machine, and the difference matters more than it sounds.
Allow about 15 minutes for a first useful comparison, plus time to get representative assembly from the code you actually care about. The examples use Ubuntu's installed LLVM 20.1.8 package and the Skylake model. You need llvm-mca-20; Clang is only needed if you want to generate assembly from C or C++.
Start by recording the version and the target models available on this machine. Version matters here because scheduling models and report details can change between LLVM releases.
llvm-mca-20 --version
On the system used for this guide, the result begins with Ubuntu LLVM version 20.1.8 and reports x86_64-pc-linux-gnu as the default target. It also lists the registered architectures. The analyser defaults to the host target and autodetects a CPU unless you select one explicitly.
Checkpoint: if the command is missing, stop here and install the package that supplies it through your normal system administration process. Nothing below needs root privileges.
Use an input file while you are learning so each run analyses exactly the same instructions. This example uses Intel syntax and keeps the kernel deliberately small.
.text
.intel_syntax noprefix
add eax, ebx
imul eax, ecx
mov edx, eax
Save those lines as kernel.s. The .intel_syntax directive tells the assembler parser which syntax to expect; without it, use the target's normal assembly syntax instead. Comments beginning with LLVM-MCA-BEGIN and LLVM-MCA-END can later restrict analysis to a region. Add no markers and the default region covers every instruction in the input.
There is no state to undo here, this step only creates a text file. If the file lives in a shared or valuable directory, write to a new path rather than overwriting an existing source file.
Choose the target triple and CPU deliberately whenever you are comparing results. The triple selects the target environment; -mcpu selects the processor scheduling model. The result below uses five simulated iterations.
llvm-mca-20 -mtriple=x86_64-pc-linux-gnu -mcpu=skylake -iterations=5 kernel.s
The output starts with values similar to:
Iterations: 5
Instructions: 15
Total Cycles: 24
Total uOps: 15
Dispatch Width: 6
uOps Per Cycle: 0.63
IPC: 0.63
Block RThroughput: 1.0
Iterations is the number of times the instruction block is simulated, so the three instructions produce 15 simulated instructions. IPC is simulated instructions divided by simulated cycles. Block RThroughput is the reciprocal of the block throughput reported by the scheduling model. These are model outputs, not promises about wall-clock performance.
The report then shows an instruction-information table: each instruction's micro-op count, latency, reciprocal throughput, and whether the model marks it as a possible load or store. The resource sections show how the model spreads pressure across execution resources such as Skylake ports.
Checkpoint: the command must exit with status zero and print Iterations. A parse error or unsupported instruction returns status one and writes an error to standard error.
Add -timeline when the summary suggests a dependency or scheduling problem. It prints a compact character timeline for each instruction and iteration.
llvm-mca-20 -mtriple=x86_64-pc-linux-gnu -mcpu=skylake -iterations=2 -timeline kernel.s
The timeline includes rows like D=eeeER. The exact columns depend on the model and the run, but the view distinguishes dispatch, execution, write-back, and retire stages. The sample also ends with average wait times, useful for showing whether an instruction spent time waiting in a scheduler queue or waiting for operands.
Keep the timeline small while investigating. Its default maximum is 80 cycles and its default iteration limit is 10. Use -timeline-max-cycles=0 for no cycle limit, or set a smaller explicit limit when a large kernel would otherwise make the report unreadable.
Change one instruction or one scheduling option at a time, then compare the same summary fields. Write the output to a new file rather than replacing a previous result:
llvm-mca-20 -mtriple=x86_64-pc-linux-gnu -mcpu=skylake -iterations=100 -o kernel-skylake.txt kernel.s
With -o, the report is written to the named file. An existing report survives if you just choose a new filename. If you genuinely need to replace one, check the path first: redirection and output-file replacement are destructive to that report, though they never touch the assembly input itself.
To compare processors, repeat the command with another valid -mcpu value. Do not let the host CPU stand in silently for a deployment target. A model may not even exist for the processor you actually want, and a model's own quality limits how useful the result can be.
llvm-mca-20 parses assembly into LLVM machine instructions and simulates a scheduling pipeline. It does not model instruction fetch, decoding, branch prediction, cache hierarchy, memory types, store-to-load forwarding, or every serialising operation. Its load and store unit is deliberately simplified and assumes loads and stores do not alias, by default.
That default is a common distraction trap. If aliasing could affect the code sequence, test the conservative alternative rather than presenting the optimistic report as fact. The installed command accepts -noalias, which explicitly assumes loads and stores do not alias; the manpage describes that as the default behaviour anyway. There is no command-line option for a richer cache or memory-consistency model.
Unsupported instructions normally stop the analysis outright. -skip-unsupported-instructions=any makes the command carry on, but silently drops unsupported instructions from the simulation. Use it only when you understand exactly what got omitted, and label the result accordingly. Fixing the input, or picking a scheduling model that supports it, beats a silently incomplete comparison every time.
The -json option prints the requested supported views as JSON, handy for a small comparison script. Not every view survives the trip: the manpage specifically warns that bottleneck-analysis output is not included in JSON.
llvm-mca-20 -mtriple=x86_64-pc-linux-gnu -mcpu=skylake -iterations=100 -json kernel.s > kernel-skylake.json
python3 -m json.tool kernel-skylake.json > /dev/null
That second command is only a syntax check. It proves the JSON parses, not that the selected CPU or the input represents your production workload.