Blog / Go

  • go
  • tcp
  • networking
  • net-listener
  • linux
  • debugging

Why a Go TCP Echo Server Stalls at 10,000 Clients, Not 10

Your echo server is twelve lines long, it handles ten test clients without complaint, and then a load test with 10,000 makes it hang. The tempting conclusion is "goroutines don't scale". They do. Ten thousand goroutines is a quiet afternoon for the Go runtime, so the stall is almost always somewhere else.

There are three usual suspects, and they look similar from the client side: the accept loop, the kernel's listen queue, and file descriptors. Here is how to tell which one you have.

Suspect one: the accept loop does the work itself

This is the classic. The handler runs inline, so the server serves one client at a time.

for {
	conn, err := ln.Accept()
	if err != nil {
		log.Fatal(err)
	}
	handle(conn) // blocks until this client disconnects
}

With ten short-lived clients you never notice, because each finishes before the next matters. Add one client that connects and sits there, and every other client queues behind it. The fix is one keyword, but the full loop is worth writing out properly, because the error branch matters too.

package main

import (
	"errors"
	"io"
	"log"
	"net"
	"time"
)

func main() {
	ln, err := net.Listen("tcp", ":7000")
	if err != nil {
		log.Fatal(err)
	}
	for {
		conn, err := ln.Accept()
		if err != nil {
			if errors.Is(err, net.ErrClosed) {
				return
			}
			log.Printf("accept: %v", err)
			time.Sleep(50 * time.Millisecond)
			continue
		}
		go func() {
			defer conn.Close()
			_, _ = io.Copy(conn, conn)
		}()
	}
}

The sleep is there on purpose. If Accept fails because you are out of file descriptors, retrying instantly burns a CPU core and floods the log. Backing off briefly is what net/http does too.

Suspect two: the accept queue overflows

Now the loop is fine, but 10,000 clients dial at the same instant. The kernel completes handshakes and parks the connections in a queue until your program calls Accept. That queue has a fixed size, and when it is full the kernel silently drops new connection attempts.

The client does not get an error. Its SYN is simply ignored, so it retransmits, typically after about one second, then again after a few more. Your load test shows a pile of dials taking 1s, 3s or 7s, while the server looks idle. That is the stall.

Seeing it

Two commands tell you whether this is happening:

  • ss -ltn 'sport = :7000': for a listening socket, Recv-Q is the current accept queue depth and Send-Q is its maximum.
  • nstat -az TcpExtListenOverflows TcpExtListenDrops: if these climb during the test, the queue overflowed.

A tiny client reproduces it. It reports any dial that took suspiciously long:

package main

import (
	"log"
	"net"
	"sync"
	"time"
)

func main() {
	var wg sync.WaitGroup
	for i := 0; i < 10000; i++ {
		wg.Add(1)
		go func() {
			defer wg.Done()
			start := time.Now()
			c, err := net.DialTimeout("tcp", "127.0.0.1:7000", 10*time.Second)
			if err != nil {
				log.Print(err)
				return
			}
			defer c.Close()
			if d := time.Since(start); d > 500*time.Millisecond {
				log.Printf("slow dial: %v", d)
			}
		}()
	}
	wg.Wait()
}

Quick detour: where the queue size comes from

Go does not ask you for a backlog. On Linux, the runtime reads /proc/sys/net/core/somaxconn when it creates the listener and uses that value. So the knob is a sysctl, not a line of Go. Older kernels defaulted it to 128; newer ones use 4096, which is still less than 10,000 simultaneous arrivals if your accept loop is momentarily busy.

Raise it with sysctl -w net.core.somaxconn=16384 and restart the server, since the value is read at listen time. The accept queue is not the only gate (there is a separate SYN queue too), so treat the counters above as the judge, not the number you set.

Suspect three: file descriptors

Every connection costs a descriptor, plus whatever else your process has open. A shell with ulimit -n 1024 will make Accept fail with "too many open files" at about the thousandth client.

Go helps here: since Go 1.19 the runtime raises the soft limit to the hard limit at start-up on Linux. That makes the old 1024 problem mostly vanish, but only up to the hard limit. Containers and some service managers set a low hard limit, so check what the process really has:

grep 'open files' /proc/$(pgrep echo)/limits
ls /proc/$(pgrep echo)/fd | wc -l

The load generator needs descriptors too. If you run it on the same box, both processes draw on the machine, and each client connection is two sockets locally. Also remember that 10,000 connections from one address to one server port use 10,000 ephemeral ports, which fits within the default range of roughly 28,000, but a second test run before the first one's sockets leave TIME_WAIT can eat into that.

What about memory?

Goroutines start with a small stack, a few KiB, so 10,000 of them is tens of megabytes at most. The part that can surprise you is buffers. io.Copy allocates a 32 KiB buffer when no fast path applies, which across 10,000 connections can add up to hundreds of megabytes of garbage-collected memory.

On Linux, a TCP to TCP copy can use splice and skip that buffer, but do not rely on it. If memory matters, measure it with a real run rather than trusting the arithmetic, and consider a smaller pooled buffer.

A quick order of checks

  1. Is the handler started with go? If not, stop there.
  2. Do nstat overflow counters rise during the test? Raise somaxconn.
  3. Does the log show "too many open files"? Check /proc/PID/limits for the hard limit.
  4. Is Accept error handling spinning? Add the back-off.

The satisfying part is that none of this is a Go problem. The runtime was ready for 10,000 clients long before the kernel queue and the process limits were.