Blog / Systems Programming

  • linux
  • mmap
  • page-cache
  • go
  • performance
  • systems-programming

A Zero-Copy mmap Read Still Pays: Faults, Copies, Stalls

"mmap is zero-copy" is true and also a bit of a fib. The bytes are not copied into your buffer, but the work does not vanish; it moves into page faults, and from there into a few places where a copy quietly reappears. I measured it rather than trusting the folklore.

Here is a small Go program that sums every byte of a file twice: once with read(2) into a 64 KiB buffer, once through mmap. It also asks the kernel how many page faults each pass caused.

func faults() (minor, major int64) {
	var ru syscall.Rusage
	if err := syscall.Getrusage(syscall.RUSAGE_SELF, &ru); err != nil {
		panic(err)
	}
	return ru.Minflt, ru.Majflt
}

// read(2) path
buf := make([]byte, 64<<10)
for {
	n, err := f.Read(buf)
	s += sum(buf[:n])
	if err != nil {
		break
	}
}

// mmap path
data, err := syscall.Mmap(int(f.Fd()), 0, int(fi.Size()),
	syscall.PROT_READ, syscall.MAP_SHARED)
if err != nil {
	panic(err)
}
defer syscall.Munmap(data)
s = sum(data)

The fault counter tells the story

I ran it on a 512 MiB file already sitting in the page cache (Linux 6.8, Go 1.22, eight cores). The output was the same on both runs:

read  minor=16    major=0 time=336ms
mmap  minor=8194  major=0 time=240ms

The read pass barely faulted at all. The mmap pass took 8194 minor faults to walk 512 MiB, and mmap was faster here, by a margin I would not bank on. The sum loop is byte-at-a-time and dominates both runs, so this is a sanity check and not a benchmark.

Those 8194 are not random. 512 MiB is 131,072 pages of 4 KiB, and 131,072 divided by 8192 is 16. So each fault mapped about 16 pages, or 64 KiB.

Quick detour: why not one fault per page?

Hang on, why 16? Because of "fault-around". When you touch a file-backed page that is already in the page cache, the kernel does not map just that one page; it maps the neighbouring cached pages in the same go. The default window is 64 KiB, which matches the arithmetic above.

I could not read the tunable on my machine (it lives in debugfs and needs root), so treat the 64 KiB as my inference from the numbers. Either way, it is why mmap is not as bad as "one fault per 4 KiB page" suggests.

Where the copy comes back

Nobody maps a file just to sum it. Usually the bytes go somewhere else, and that is where copies reappear:

  • Writing to a socket. write(2) from a mapped region still copies from the page cache into the socket buffer. You saved the copy into user space and then paid one out again.
  • Writing to a private mapping. With MAP_PRIVATE, the first write to a page makes the kernel copy the whole page first (copy-on-write).
  • Parsing. Decompressing, decoding or building structures from the bytes copies them into new memory anyway.

For file-to-socket specifically, sendfile(2) skips user space entirely, which mmap cannot do. The one path that avoids the CPU copy from the page cache altogether is O_DIRECT, where the device DMAs straight into your buffer. It needs aligned buffers and offsets, and you lose the cache, so it is a database engine's tool, not a casual one.

Under load, the costs stop being uniform

My test file was warm. Under load, it often is not, and that is where the two approaches diverge. A cold page turns a cheap minor fault into a major fault, which means waiting for the disk.

With read(2) the wait happens inside a syscall you can see in strace, and you can push it onto a worker thread. With mmap the wait happens inside an ordinary memory load. There is no syscall to intercept, and nothing in an event loop can make it asynchronous; the whole thread just stops.

Other things get worse when many threads are involved:

  • Mapping and unmapping change the process address space, and munmap has to tell every CPU that might have cached the old translations (a TLB shootdown). Do that often on a many-core box and it shows up.
  • Reclaim under memory pressure evicts mapped pages, and the next access faults them straight back in. That thrash is hard to see from user space.
  • Faults contend on kernel address-space locks. Recent kernels have narrowed this with per-VMA locking, but I would still measure on your own kernel.

The CIDR 2022 paper "Are You Sure You Want to Use MMAP in Your Database Management System?" (Crotty, Leis and Pavlo) goes through these problems in database engines and is worth a read if you are considering mmap for anything serious.

The failure mode read(2) does not have

An I/O error during read(2) gives you an error return. During a mapped load it gives you SIGBUS. The same happens if another process truncates the file while you have it mapped.

In Go, that is a fatal "unexpected fault address" crash unless you opt in with debug.SetPanicOnFault(true) on that goroutine and recover. In Rust, this is precisely why memmap2::Mmap::map is unsafe: the compiler cannot promise that nobody else will change the file underneath your &[u8].

What I would actually do

Start with read(2) and a decent buffer. It is predictable, its errors are errors, and the copy is cheap compared with most of what you do next. Reach for mmap when you have measured it and the workload is read-mostly, random access into a large file that mostly stays cached.

Whichever you pick, watch the counters. getrusage, /usr/bin/time -v or perf stat -e minor-faults,major-faults will show you in seconds whether your "zero-copy" read is one big bulk of quiet work or eight thousand small interruptions.