Blog / Linux

  • linux
  • inodes
  • filesystems
  • hard-links
  • unlink
  • systems-programming

Why Deleting a Running Binary Doesn't Crash It

You can rm a program while it is running and it carries on happily. Try to overwrite the same program with cp, though, and you get "Text file busy". Deleting is the more destructive act, so that looks backwards. It isn't, and the reason is that "delete" is not what Unix actually does.

Reproduce it in thirty seconds

Copy a harmless binary somewhere, start it, then remove it:

cp "$(command -v sleep)" /tmp/demo
/tmp/demo 300 &
ls -li /tmp/demo
rm /tmp/demo
ls -l /proc/$!/exe

The first ls -li shows an inode number, then a link count of 1 in the third column. After the rm, the process is still alive and /proc/PID/exe reports /tmp/demo (deleted). The program did not notice anything happen.

A filename is not the file

A directory is just a table mapping names to inode numbers. The inode holds everything else: size, owner, permissions, timestamps and where the data blocks live. The name is a pointer to the file, not the file itself.

That is why rm is implemented by a system call named unlink. It removes one directory entry and decrements the inode's link count. It says nothing about the data.

The kernel frees an inode and its blocks only when two things are true:

  • the link count is zero (no names point at it), and
  • nothing in the kernel still holds a reference to it.

Who holds the reference

A running executable is held by the kernel in several ways. The binary is memory-mapped into the process, so the mapping keeps the inode alive. The kernel also keeps the exec'd file open for /proc/PID/exe. Open file descriptors, the current working directory and any other mmaps count too.

So after rm, the link count is 0 but the in-kernel reference count is not. The inode stays until the last holder goes away, at which point the blocks are released. Pages not yet faulted in are still read from the same inode, which is why the process doesn't fall over later either.

Detour: link counts you can see

Quick detour, because this is easy to poke at. Make a second name for the same inode:

echo hello > a
ln a b
stat -c '%i %h %n' a b
rm a
cat b

Both names print the same inode number and a link count of 2. After rm a, b still has the contents, because the count went from 2 to 1. Nothing was "copied"; there was only ever one file. (Hard links can't cross filesystems for the obvious reason: an inode number only means something within its own filesystem.)

So why does overwriting fail?

Because overwriting changes the inode's data, and the process is executing from that data. If you could rewrite the bytes under a running program, its next page fault would load instructions from the middle of your new file. That ends badly.

Linux blocks it: opening a file for writing while it is being executed fails with ETXTBSY.

$ cp /bin/ls /tmp/demo
cp: cannot create regular file '/tmp/demo': Text file busy

The same happens with echo x >> /tmp/demo. Unlinking changes the directory, not the inode's contents, so it is allowed. The kernel protects the data, not the name.

What this means for upgrades

This is why well-behaved installers write the new version to a temporary name in the same directory and then rename it over the old one. The rename atomically swaps the name to point at a new inode. Running processes keep the old inode; anything started afterwards gets the new one.

  • rm then cp: works, with a brief window where the file doesn't exist.
  • cp straight over it: fails with "Text file busy" for a running binary.
  • Write to a temp file, then mv over it: works and never leaves a missing file.

The catch is that old and new versions now run side by side until the old processes restart. That is exactly why services need a restart after a package upgrade, and why tools like needrestart exist: they look for processes whose mapped files are marked deleted.

A sharper edge: in-place writes to files that are merely mmapped (data files, and in some situations shared libraries) are not always blocked by ETXTBSY. A process reading changed pages mid-flight can misbehave or crash. Replacing by rename avoids the whole class of problem.

Getting the file back

If you deleted a binary that is still running, the data is intact. Copy it out through procfs:

cp /proc/PID/exe /tmp/recovered
chmod +x /tmp/recovered

That reads the inode through the process's own reference and writes it under a fresh name. Note this gives you the file's contents, not its original path, so you will need to put it back yourself and check the permissions.

The disk-space side of this (a deleted log held open by a daemon, quietly eating the filesystem) is the same mechanism seen from the other end, and the blocks only come back when the last holder exits.