Serialise Shell Jobs Safely with flock

flock stops two copies of the same cron job or script from running at once, using an ordinary file as the lock. By the end you will have one job that waits its turn, a second invocation that can fail fast when required, and you will know exactly which exit status means the lock was unavailable. The examples use the flock command from util-linux.

You need a Linux system with util-linux installed and a POSIX-compatible shell. The local reference here is util-linux 2.39.3; the first flock found on some PATH values can be a different build, so use /usr/bin/flock when comparing output on this machine. Allow about 10 minutes, and no elevated privilege is needed while the lock file lives in a directory you can already write to.

1. Confirm the command and choose a lock path

Check which implementation you are actually about to use:

/usr/bin/flock --version
dpkg-query -W -f='${Package} ${Version}\n' util-linux

On the reference system, the package reports util-linux 2.39.3, and the command normally reports its version and exits successfully. If /usr/bin/flock is absent, run command -v flock and check that installation's --version output before trusting anything version-specific below.

Pick a lock file every competing invocation can reach, but that ordinary users cannot overwrite casually. For a per-user job, $HOME/.cache/example-job.lock is reasonable; for the self-contained checks below, use a temporary directory instead:

lock_dir=$(mktemp -d)
lock_file="$lock_dir/example-job.lock"
trap 'rmdir "$lock_dir" 2>/dev/null || :' EXIT

The file does not need to contain data; flock creates a file or directory at that path when permissions allow it. The lock is attached to the open file, not to its text, so deleting and recreating the lock path while another process is using it can create two unrelated locks. Do not put this file somewhere another process cleans up while the job is running.

2. Make a job wait for the lock

Wrap the command that must not overlap:

/usr/bin/flock "$lock_file" -c 'printf "job started\n"; sleep 2; printf "job finished\n"'

The default is an exclusive lock. If another process already owns it, this invocation waits until that process closes the locked file. The exit status of a successful child is passed straight back, so your usual failure handling around the real job still applies.

Run two copies close together to see the serialisation itself:

/usr/bin/flock "$lock_file" -c 'printf "first\n"; sleep 2' &
first_pid=$!
/usr/bin/flock --verbose "$lock_file" -c 'printf "second\n"'
wait "$first_pid"

The second invocation only prints once the first lock is released. With --verbose, flock also reports how long it took to acquire the lock, though exact wording varies with the util-linux version; the ordering and a successful exit status are the useful checks.

3. Refuse overlap instead of waiting

Add --nonblock when a missed schedule beats a delayed run:

/usr/bin/flock "$lock_file" -c 'sleep 3' &
holder_pid=$!
if /usr/bin/flock --nonblock --conflict-exit-code 75 "$lock_file" -c 'printf "unexpected work\n"'; then
    printf 'lock acquired\n'
else
    status=$?
    printf 'lock unavailable, status %s\n' "$status"
fi
wait "$holder_pid"

The second command does not wait; it should print lock unavailable, status 75. --conflict-exit-code must be between 0 and 255 and defaults to 1, but picking a distinct value lets a wrapper tell lock contention apart from a genuine error in the protected command.

When a short delay is acceptable instead of an outright refusal, bound the wait:

/usr/bin/flock "$lock_file" -c 'sleep 3' &
holder_pid=$!
/usr/bin/flock --wait 0.5 --conflict-exit-code 75 "$lock_file" -c 'printf "work\n"'
status=$?
printf 'bounded attempt status: %s\n' "$status"
wait "$holder_pid"

Fractional seconds are allowed. If the lock is still held when the timeout expires, flock returns the conflict status and never runs the protected command at all. A timeout of zero behaves the same as non-blocking mode.

4. Use a file descriptor inside a script

For a script with several commands, opening one descriptor keeps the lock tied to the shell process itself, instead of spawning a separate flock wrapper for every command:

#!/bin/sh
lock_file="${HOME}/.cache/example-job.lock"
mkdir -p "${HOME}/.cache" || exit 1
exec 9>"$lock_file" || exit 1
if ! /usr/bin/flock --nonblock 9; then
    printf '%s\n' 'another example job is already running' >&2
    exit 75
fi
printf '%s\n' 'doing protected work'
# Commands here run while descriptor 9 remains open.
printf '%s\n' 'protected work complete'

The > redirection creates the file if needed and requires write permission; the file's mode is not what provides the lock. The descriptor stays open until the shell exits or closes it, so watch out for background work that must not inherit it.

Safety boundaries and recovery

Safety boundary: do not use rm to clear a lock while a job might still be running. Removing the pathname does not release the old open-file lock, and it can let a new pathname acquire an unrelated lock instead. Stop the holder cleanly, then leave the empty lock file in place; the next invocation acquires it normally. A lock releases automatically when its file descriptor closes, including when the protected command exits.

flock does not detect deadlocks. Two jobs can still wait on locks in an order that never completes, so keep your locking order consistent everywhere. Locks on NFS or CIFS have limited support and can fail depending on mount options; if a command works locally but not over a network mount, check the filesystem and mount behaviour before you weaken the coordination to compensate.

The --close option closes the locking descriptor before the wrapped command starts, for when a child process must not inherit the lock. --no-fork replaces flock with the command while keeping the lock, and is incompatible with --close; neither is needed for the basic wrapper or descriptor patterns above.

Done means