lockfile stops two cron jobs writing the same report file at once, the kind of race that leaves you with half a CSV and a support ticket. It lets only one cooperating process touch a shared resource at a time, then releases it reliably. The examples use lockfile v3.24 from procmail package 3.24-1ubuntu2.
Allow about fifteen minutes. You need a shell, the installed lockfile command and write access to a working directory. The normal examples are unprivileged; do not use sudo unless the directory or resource genuinely belongs to a privileged service account.
Check the binary and version before copying an option into a script. This is read-only:
$ command -v lockfile
/usr/bin/lockfile
$ lockfile -v
lockfile v3.24 2022/03/02
$ dpkg-query -W -f='${Package} ${Version}\n' procmail
procmail 3.24-1ubuntu2
The command creates semaphore files. A lock is not a kernel-wide guarantee for every program that opens the protected file: it works only when all competing scripts follow the same convention. A program that ignores the semaphore can still access the resource.
Use a private directory so the test cannot collide with another user. The lock file is deliberately named with the conventional .lock suffix:
$ workdir=$(mktemp -d)
$ lockfile "$workdir/report.lock"
$ printf 'lock status: %s\n' "$?"
lock status: 0
$ ls -l "$workdir/report.lock"
-r--r--r-- 1 ... report.lock
A status of zero means the file was created. The installed program makes lock files read-only, so ordinary removal can fail. Keep the directory path in workdir; it is the only test state you need to clean up.
Checkpoint: confirm that the lock exists before continuing.
$ test -f "$workdir/report.lock" && echo 'lock exists'
lock exists
By default, lockfile waits eight seconds between attempts and retries forever. That suits a long-running batch job, but it is also an easy way to make a test appear frozen. Use a short wait and a finite retry count when you need a bounded operation:
$ lockfile -1 -r 0 "$workdir/report.lock"
Sorry, retries limit reached
$ printf 'exit status: %s\n' "$?"
exit status: 1
Here -1 means wait one second between attempts, and -r 0 permits no retry after the first failed creation. The exact diagnostic text comes from this procmail release, but scripts should test the exit status rather than parse that text.
Do not rely on the default infinite retry behaviour in an unattended script unless waiting forever is an explicit requirement. A stale lock or a crashed peer otherwise leaves the process waiting without a useful deadline.
Acquire the lock before touching the shared resource, and remove it after the work succeeds or fails. An EXIT trap gives the shell a cleanup path for ordinary errors and signals:
#!/bin/sh
set -eu
target=/path/to/shared-report
lock="$target.lock"
lockfile -1 -r 30 "$lock" || {
printf '%s\n' "another process still owns $lock" >&2
exit 1
}
trap 'rm -f -- "$lock"' EXIT HUP INT TERM
# Replace this with the short operation that must be serialised.
printf '%s\n' 'update the shared report here'
exit 0
The placeholder path is not safe to run until you replace it with a real path you control. The trap uses rm -f because the lock is read-only, and -- so a lock name beginning with a hyphen cannot be interpreted as an option by rm.
Checkpoint: while the script is paused inside its critical section, another copy should return non-zero after its retry limit. Test this with a harmless command before placing database, mail or file replacement work inside the section.
Remove the test lock explicitly, then remove the empty temporary directory. These are state-changing commands, so check the variable before using it:
$ test -n "${workdir:-}" && test -d "$workdir"
$ rm -f -- "$workdir/report.lock"
$ rmdir -- "$workdir"
$ test ! -e "$workdir" && echo 'test state removed'
test state removed
If your real critical section exits early, the EXIT trap performs the equivalent lock removal. If a process is killed abruptly, inspect the lock before forcing its removal: the timeout option exists for abandoned locks, but it is not proof the original owner has stopped running.
Use -l to set a lock timeout in seconds. Once the lock has not been modified for that period, lockfile may remove it and retry. After a forced removal it waits for the suspend interval, sixteen seconds by default, to reduce the risk of immediately deleting a newly created lock:
$ lockfile -1 -r 2 -l 3600 "$workdir/report.lock"
$ printf 'exit status: %s\n' "$?"
exit status: 0
That example only suits a case where an hour-old lock cannot represent live work in your application. Choosing the timeout needs knowledge of the longest legitimate run and the consequences of two processes overlapping: too short breaks a healthy job, too long delays recovery.
Warning: forcing a stale lock is a coordination decision, not routine housekeeping. Check the owning process, service logs and lock path first. Never delete another program's lock merely because it is inconvenient.
The -! option inverts the exit status when the lock already exists, which can be useful as the condition of a shell loop. Its behaviour is deliberately narrow: failures for other reasons remain failures. The manual warns that the result is not intuitive, so a direct status test is usually clearer for a single operation:
$ while lockfile -1 -r 1 -! "$workdir/report.lock"; do
> printf '%s\n' 'lock already exists; retrying'
> sleep 1
> done
Do not use this example without understanding which status your loop body expects. For a normal script, acquire once, handle a non-zero result, and keep cleanup in the EXIT trap.
- must be written with a ./ prefix.lockfile removes the files it created during that invocation.-ml locks the system mailbox and -mu unlocks it when the spool permissions or setgid installation allow it. Do not experiment with those options on a production mailbox; they are not a substitute for the ordinary application lock pattern.rm -f.