Why Your Daemon Keeps a Zombie Child Until You Call wait()
You run ps and find a handful of processes marked Z and <defunct>. They use no CPU and no real memory, and kill -9 does nothing to them. They are already dead. They stay in the process table because the parent has not asked how they died.
Reproducing it in a few lines of Go
This program starts a child that exits immediately, then looks at its state in /proc before and after calling Wait.
package main
import (
"fmt"
"os"
"os/exec"
"strings"
"time"
)
// state returns the one-letter state from /proc/PID/stat.
func state(pid int) string {
b, err := os.ReadFile(fmt.Sprintf("/proc/%d/stat", pid))
if err != nil {
return "gone"
}
s := string(b)
return string(s[strings.LastIndexByte(s, ')')+2])
}
func main() {
cmd := exec.Command("true")
if err := cmd.Start(); err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
time.Sleep(100 * time.Millisecond)
fmt.Println("before Wait:", state(cmd.Process.Pid))
if err := cmd.Wait(); err != nil {
fmt.Fprintln(os.Stderr, err)
}
fmt.Println("after Wait: ", state(cmd.Process.Pid))
}
On Linux this prints Z and then gone. The child finished within microseconds, yet it sat there for the whole sleep. Nothing about the child changed between the two lines; only the parent's behaviour did.
What the kernel is actually holding on to
When a process exits, the kernel frees its memory, closes its file descriptors and drops its address space. It keeps a small record: the PID, the exit status and some resource usage figures. That record exists so the parent can collect it with wait(), waitpid() or waitid().
Once the parent has read it, the record is released and the PID becomes available again. Until then, the entry is the zombie. It is a receipt nobody has signed for.
Quick detour: why not just discard it?
Hang on, why does the kernel not simply throw the status away? Because it cannot know whether you want it. A shell needs it for $?, a supervisor needs it to decide whether to restart, and a build tool needs the rusage numbers. Nothing tells the kernel in advance that you do not care.
There is an opt-out, though. If the parent sets SIGCHLD to SIG_IGN (or uses SA_NOCLDWAIT), children are reaped automatically. The catch is that a later wait() then blocks until every child has gone and fails with ECHILD, so you lose the status for good.
Why it matters for a long-running daemon
One zombie is harmless. A daemon that forks a child per request and never waits leaks one per request. Each costs a process table slot and a PID, and the limits are real:
- the system-wide ceiling in
/proc/sys/kernel/pid_max - per-user process limits (
RLIMIT_NPROC) - the
pids.maxlimit on a cgroup, which is what containers usually hit first
When one is exhausted, fork fails with EAGAIN and the daemon looks mysteriously broken while using almost no memory.
Fixing it in your own code
If you started the child, you wait for it. In Go that means every successful cmd.Start() needs a matching cmd.Wait(), even when you do not care about the result. For fire-and-forget, reap in a goroutine:
if err := cmd.Start(); err != nil {
return err
}
go func() {
if err := cmd.Wait(); err != nil {
log.Printf("child %d: %v", cmd.Process.Pid, err)
}
}()
Note that cmd.Run() and cmd.Output() already call Wait. The leak only shows up when you split Start from Wait and an early return skips the second half.
The orphans you never started
The nastier case is a grandchild. If a process dies while its children live on, the kernel reparents them to PID 1, or to the nearest ancestor that declared itself a subreaper. When those orphans exit, whoever inherited them has to wait.
On a normal host, init or systemd does this without being asked. In a container, PID 1 is whatever you set as the entrypoint, and a Go binary does not reap anything it did not start. Orphans pile up as zombies until the cgroup runs dry.
Options, from least to most effort:
- Run the container with an init process, for example
docker run --init, or usetinias the entrypoint. - Do not spawn processes you do not track; where you must, keep the
Waitgoroutine above. - Make your daemon a subreaper with
prctl(PR_SET_CHILD_SUBREAPER)and loop onwait4(-1, ...)yourself.
The third is the one with sharp edges. A blanket wait4(-1) competes with os/exec: if your loop steals a child's status, the matching cmd.Wait() returns an error because the process is already gone. Reap only PIDs you do not own, or put the reaper in a separate small process.
Checking whether you have a problem
ps -eo pid,ppid,stat,comm | awk '$3 ~ /^Z/'
The ppid column is the useful one. It names the process that is failing to reap. Killing the zombie does nothing, but killing or fixing that parent does: once the parent dies, the zombies are reparented to PID 1, which (if it behaves) clears them straight away.