Home / Alt manpages / llvm-mca-18(1)

  • llvm-mca-18(1)
  • User command
  • linux

Measure Assembly Throughput with llvm-mca-18

You will run a small assembly block through llvm-mca-18, read its throughput and resource-pressure report, then narrow the analysis to a named region and a chosen number of iterations. The installed command is Ubuntu LLVM 18.1.3 from package version 1:18.1.3-1ubuntu1.

Allow about 15 minutes. You need a shell, a readable assembly file and an LLVM target with a scheduling model. These examples only read input and write reports. They do not assemble, install, alter a service or require sudo.

1. Check the installed analyser

Confirm which executable will run and record its version:

$ command -v llvm-mca-18
/usr/bin/llvm-mca-18
$ llvm-mca-18 --version
Ubuntu LLVM version 18.1.3
  Optimized build.

The version output also reports the default target and host CPU. On this machine the default target is x86_64-pc-linux-gnu and the host CPU is skylake. The host CPU is only a default: use -mcpu when you want a reproducible model that is not tied to the machine running the command.

Checkpoint

Keep the version and CPU model with any performance result. A report made with a different scheduling model is not automatically comparable.

2. Create a small analysis region

Put the instructions you want to study in an assembly file. The special comments mark a region; the surrounding comments are ignored by the analyser:

# LLVM-MCA-BEGIN add-loop
addl %esi, %edi
addl %edx, %edi
# LLVM-MCA-END

Save that as /tmp/mca-demo.s for a disposable test, or use a project file that you already trust. The markers are comments in this x86 AT&T syntax, so they do not change the instructions being measured. A region may have a name after LLVM-MCA-BEGIN; the ending marker does not repeat it.

Run the default report with an explicit CPU:

$ llvm-mca-18 -mcpu=skylake /tmp/mca-demo.s

[0] Code Region - add-loop

Iterations:        100
Instructions:      200
Total Cycles:      203
Total uOps:        200

Dispatch Width:    6
uOps Per Cycle:    0.99
IPC:               0.99
Block RThroughput: 0.5

The rest of the report contains instruction information and resource pressure. The default simulation runs 100 iterations, so two instructions produce 200 counted instructions. IPC is instructions per cycle for the simulation. Block RThroughput is the model's steady-state throughput estimate for one block, expressed in cycles per iteration. Do not read either number as a benchmark of a real running process.

3. Read the resource and instruction sections

In the instruction information view, the columns show micro-operations, latency and reciprocal throughput, followed by load, store and side-effect flags. For the two addl instructions above, the installed Skylake model reports one micro-operation and latency 1 for each.

The resource section names modelled execution resources such as SKLPort0 and SKLPort1. Resource pressure is reported per iteration and by instruction. A high value on one resource can explain why a sequence does not scale as well as its instruction count suggests. The resource view is enabled by default; -resource-pressure makes that choice explicit.

Use -instruction-tables when you want static resource information from the processor model without simulating the code. That is a different question from the normal resource-pressure view. Use -show-encoding if the instruction information needs machine-code encodings as well.

4. Make the simulation and timeline useful

Set the iteration count when comparing reports or when a short report is easier to inspect:

$ llvm-mca-18 -mcpu=skylake -iterations=2 -timeline /tmp/mca-demo.s

Timeline view:
Index     0123456

[0,0]     DeER ..   addl  %esi, %edi
[0,1]     D=eER..   addl  %edx, %edi

The real output includes every selected iteration. The timeline letters describe pipeline stages and waiting; the exact cycles depend on the target model and instruction dependencies. The default timeline is limited to 10 iterations and 80 cycles. Change those limits with -timeline-max-iterations and -timeline-max-cycles; use a cycle value of zero for no cycle limit.

For machine-readable processing, request JSON:

$ llvm-mca-18 -mcpu=skylake -iterations=1 -json /tmp/mca-demo.s > /tmp/mca-demo.json
$ head -n 8 /tmp/mca-demo.json
{
  "CodeRegions": [
    {
      "InstructionInfoView": {

The JSON contains named code regions, instruction data, summary values and target information. Not every optional view is supported in JSON: the manual specifically excludes bottleneck-analysis output. Treat a successful parse and the recorded simulation parameters as part of your verification.

5. Compare CPUs deliberately

Choose a CPU model rather than relying on host autodetection when the result will be reviewed elsewhere:

$ llvm-mca-18 -mcpu=skylake -iterations=100 /tmp/mca-demo.s > /tmp/skylake.txt
$ llvm-mca-18 -mcpu=znver3 -iterations=100 /tmp/mca-demo.s > /tmp/znver3.txt
$ grep -E '^(Iterations|Total Cycles|IPC:|Block RThroughput)' /tmp/skylake.txt /tmp/znver3.txt

Only compare the values after checking that the assembly syntax, target architecture, CPU, iteration count and selected regions are the same. -march selects an architecture, while -mcpu selects a processor model within the target. If the input is not for the target you selected, a successful report can still be the wrong report.

For compiler output, the documented workflow is to pipe assembly into the analyser. Keep the compiler target and assembler syntax aligned with the analyser:

$ clang demo.c -O2 --target=x86_64 -S -o - | llvm-mca-18 -mcpu=skylake
$ clang demo.c -O2 --target=x86_64 -masm=intel -S -o - | llvm-mca-18 -mcpu=skylake

Intel syntax is detected from an .intel_syntax directive at the start of the input. The report's output syntax normally follows the input. Verify the emitted assembly before treating the report as an analysis of the code you intended.

6. Diagnose errors without changing the machine

A successful exit status is zero. On an error, llvm-mca-18 writes a message to standard error and returns 1. Capture both streams when investigating a script:

$ llvm-mca-18 -mcpu=skylake /path/to/input.s > /tmp/mca-report.txt 2> /tmp/mca-error.txt
$ status=$?
$ printf 'status=%s\n' "$status"
$ test "$status" -eq 0 && sed -n '1,25p' /tmp/mca-report.txt || cat /tmp/mca-error.txt

Common traps are an absent input file, assembly for a different architecture, an unsupported instruction, or a CPU name that this build does not recognise. Check the file path and ask the command for its option summary with -help. Do not "fix" a model mismatch by changing the source instruction sequence until you have established which target and syntax the input uses.

The tool is a static model, not a hardware measurement. Its accuracy depends on LLVM's scheduling model for the chosen processor. Memory behaviour, operating-system activity, frequency changes and front-end effects outside that model can make a real benchmark differ. Use llvm-mca to explain likely bottlenecks, then measure a complete program separately when the question is end-to-end performance.

Done means

  • The installed LLVM version and chosen CPU model are recorded.
  • The input contains the intended assembly and, where useful, named LLVM-MCA-BEGIN and LLVM-MCA-END markers.
  • The report's iteration count, IPC and block throughput have been read in the context of the selected model.
  • Resource pressure or timeline output was enabled only when it answers the question being investigated.
  • Any JSON or redirected report has been checked for a zero exit status and retained with its target parameters.
  • No privileged command or system-changing action was needed.