epoll Said Readable, Then read() Blocked: Four Ways That Happens
You call epoll_wait, it tells you a socket is readable, you call read, and your whole event loop freezes. The kernel did not lie to you, exactly. It told you something was true a moment ago, and you treated it as a promise.
The rule that fixes this is short: readiness is a hint, so every descriptor you hand to epoll must be non-blocking. The more interesting part is why the hint can go stale, so let's go through the cases.
What epoll actually reports
An event means "at the moment I checked, an operation would not have blocked". Between that check and your next syscall, anything can happen. Nothing holds a lock on the socket for you.
The man pages say this outright. They recommend non-blocking descriptors alongside select, poll and epoll for exactly this reason, and the select(2) page lists the failure under its bugs section. It is documented behaviour, not a kernel oddity.
Four ways the hint goes stale
- Another reader got there first. Two threads wait on one epoll instance, or a forked child shares the descriptor. Both wake, one reads the data, the other blocks.
- The data was discarded. The
select(2)page describes a datagram that arrives, is reported as readable, then fails its checksum and is dropped. Newer kernels check for this in the UDP poll path, but the man page's advice stands: do not rely on it. - The connection vanished before accept(). The
accept(2)page notes that a pending connection may be removed by an asynchronous network error, or by another thread, before you call it. - Writable does not mean "writable by this much". A blocking
writeof a large buffer waits until all of it is queued, even if epoll only promised room for a little.
Cases one and three are the common ones in real servers. Case four is the sneaky one, because it passes every test with small payloads.
Watching it happen
Here is a tiny reproduction in Go using only syscall. It stands in for the race: one byte arrives, epoll reports it, and a "rival reader" takes it before we do.
package main
import (
"errors"
"fmt"
"log"
"syscall"
)
func must(err error) {
if err != nil {
log.Fatal(err)
}
}
func main() {
pair, err := syscall.Socketpair(syscall.AF_UNIX,
syscall.SOCK_STREAM|syscall.SOCK_NONBLOCK, 0)
must(err)
a, b := pair[0], pair[1]
ep, err := syscall.EpollCreate1(0)
must(err)
ev := syscall.EpollEvent{Events: syscall.EPOLLIN, Fd: int32(b)}
must(syscall.EpollCtl(ep, syscall.EPOLL_CTL_ADD, b, &ev))
_, err = syscall.Write(a, []byte("x"))
must(err)
events := make([]syscall.EpollEvent, 8)
n, err := syscall.EpollWait(ep, events, -1)
must(err)
fmt.Println("epoll reports ready:", n)
buf := make([]byte, 16)
_, err = syscall.Read(b, buf) // the rival reader wins
must(err)
_, err = syscall.Read(b, buf) // our read, after the event
if errors.Is(err, syscall.EAGAIN) {
fmt.Println("EAGAIN: ready was a hint, not a promise")
}
}
Drop SOCK_NONBLOCK from the Socketpair call and the second read never returns. Same event, same data, but now the thread is parked and every other connection on that loop waits behind it. (Error handling for EINTR on EpollWait is left out to keep the example short.)
Quick detour: closing a descriptor doesn't always leave epoll
This one surprised me when I first read it. epoll does not watch file descriptor numbers, it watches the open file description underneath. If that description has been duplicated with dup or inherited through fork, then closing your descriptor does not remove it from the epoll set.
So you can get events for a socket you believe you closed. The epoll(7) page covers this, and the advice is to call EPOLL_CTL_DEL before closing. Back to the main thread.
The fix is boring on purpose
Set O_NONBLOCK on every socket, including the listener, and treat EAGAIN as a normal answer rather than an error. Then the stale hint costs you one wasted syscall instead of a stuck loop.
- For
accept: loop untilEAGAIN, and ignore transient errors such as an aborted connection. - For
write: send what fits, keep the rest, and register forEPOLLOUTonly while you have a backlog. - For multiple threads: use
EPOLLONESHOTand re-arm after handling, or give each thread its own epoll instance.
That last point also matters for the thundering herd. Many workers waiting on one listening socket all wake for one connection; EPOLLEXCLUSIVE exists to reduce that, though it does not remove the need for non-blocking accept.
Edge-triggered is the opposite trap
With EPOLLET the failure flips. You are told once when data arrives, and if you read only part of it, you may never be told again. The rule there is to read until EAGAIN, every time. Edge-triggered mode requires non-blocking descriptors anyway, since the only way to know you have drained a socket is to be told EAGAIN.
Level-triggered with non-blocking sockets is forgiving and hard to get wrong. Start there, and move to edge-triggered only when a profile gives you a reason.
Go programmers get all of this for free: the runtime's netpoller sets sockets non-blocking and parks only the goroutine on EAGAIN. It is worth knowing what it is quietly doing for you.