A cron job process overlap (or stampede) occurs when a scheduled task’s execution duration ($T_{\text{exec}}$) exceeds its schedule interval ($T_{\text{interval}}$).
When this happens, the cron daemon launches a second, third, or fourth instance of the same script concurrently. These competing processes battle for database write locks, saturate disk I/O, exhaust RAM, and can trigger catastrophic server Out-Of-Memory (OOM) crashes.
Linux flock(2) provides kernel-managed advisory locking to guarantee mutual exclusion: only one instance of your script can execute at any given time, while overlapping invocations exit cleanly or wait safely.
30-Second Locking Comparison Table
| Locking Mechanism | Kernel Primitive | Crash Safety (kill -9 / OOM) | Stale Lock Hazard | Production Verdict |
|---|---|---|---|---|
touch /tmp/job.lock | None (Userspace file check) | ❌ No | ⚠️ High (Orphan file blocks all future runs) | Unsafe (Vulnerable to TOCTOU races) |
PID File + kill -0 | Process table existence | ❌ No | ⚠️ High (False positives on PID recycling) | Legacy Only (Brittle in containers) |
Atomic mkdir Lock | VFS directory creation | ⚠️ Partial | ⚠️ High (Stale directory after power loss) | Suboptimal (Requires manual trap cleanup) |
fcntl(2) Record Lock | Kernel byte-range lock | ✅ Yes | ❌ None | Good (Standard for C/C++ daemons) |
flock(1) CLI Wrapper | Kernel advisory file lock | ✅ Yes | ❌ None | Recommended for Crontab CLI |
Embedded flock in Bash | Kernel descriptor lock (exec 200>) | ✅ Yes | ❌ None | Recommended for Production Shell Scripts |
# The Gold-Standard 1-Line Crontab Lock (Non-blocking)
*/5 * * * * /usr/bin/flock -n /var/lock/backup.lock -c '/usr/local/bin/backup-database.sh'
1. How Linux flock Works Under the Hood
Unlike homemade lockfiles that rely on the existence of a file path on disk, Linux flock(2) associates lock state directly with an open file description inside the Linux kernel VFS layer.
How Linux Kernel flock Operates
USER SPACE
┌────────────────────────────────────────────────────────────┐
│ cron Daemon (PID 1) │
│ │ │
│ └── Launches /usr/bin/flock │
│ │ │
│ ▼ │
│ flock(2) System Call (LOCK_EX | LOCK_NB) │
│ │ │
│ ▼ │
│ File Descriptor (e.g. FD 200) │
└────────────┼───────────────────────────────────────────────┘
│
▼
KERNEL VFS LAYER
┌────────────────────────────────────────────────────────────┐
│ File Descriptor Table (Current Process) │
│ │ │
│ ▼ │
│ struct file (Open File Description) │
│ │ │
│ ▼ │
│ struct inode (In-Memory Filesystem Node) │
│ │ │
│ ▼ │
│ struct file_lock (Kernel Advisory Lock Chain) │
│ │ │
│ ├── Granted: Process executes critical section │
│ └── Denied: flock -n exits immediately with code 1 │
└────────────────────────────────────────────────────────────┘
Why flock Is 100% Crash-Safe:
- When your script executes, it opens
/var/lock/myjob.lockand requests an exclusive lock (LOCK_EX). - The Linux kernel assigns the lock to the process's open file descriptor.
- If the script crashes, runs out of memory (OOM Killer), receives
kill -9, or suffers a kernel panic, the Linux kernel automatically closes all open file descriptors during process teardown. - The instant the file descriptor closes, the kernel immediately releases the lock.
- The lockfile (
/var/lock/myjob.lock) remains on disk, but the kernel lock is gone. The next scheduled cron run acquires the lock without any manual cleanup required.
2. The 4 Production flock Implementation Patterns
Pattern 1: Non-Blocking Execution in Crontab (Drop Overlaps)
If a previous execution is still running and you want the new invocation to exit immediately rather than pile up:
# Run every 5 minutes. If already running, exit immediately with status 1
*/5 * * * * /usr/bin/flock -n /var/lock/sync-orders.lock -c '/usr/local/bin/sync-orders.sh'
-n(Non-blocking): If another process holds the lock,flockreturns immediately without waiting.- If lock acquired: Runs
sync-orders.shand exits with the script's exit code. - If lock busy: Exits immediately with return code
1.
Pattern 2: Bounded Timeout Waiting (Wait for Completion)
If missing an execution is unacceptable, but waiting indefinitely creates a dangerous backlog, use a bounded timeout:
# Wait up to 60 seconds to acquire the lock before giving up
*/10 * * * * /usr/bin/flock -w 60 /var/lock/generate-invoices.lock -c '/usr/local/bin/generate-invoices.sh'
-w 60(Timeout in seconds): If another worker holds the lock,flockwill wait up to 60 seconds for the running job to finish. If the lock frees up within 60s, it acquires the lock and runs. If 60s expires, it aborts.
⚠️ Warning — The Unbounded Queue Anti-Pattern:
Never runflockin blocking mode without-nor-win high-frequency cron jobs (e.g.* * * * * flock /var/lock/job.lock ...). If the script takes 5 minutes, cron will spawn 5 waiting processes stacked behind the lock, transforming your scheduler into an unconstrained process stampede.
Pattern 3: Custom Exit Codes to Suppress Cron Email Alerts
By default, when flock -n fails to acquire a lock, it exits with return code 1. If your crontab sends email alerts for non-zero exit codes, lock contention triggers false-alarm emails.
Use -E <code|0> to customize the exit code on lock contention:
# Return exit code 0 when skipped due to lock contention (suppresses cron mail)
*/5 * * * * /usr/bin/flock -E 0 -n /var/lock/cache-prune.lock -c '/usr/local/bin/cache-prune.sh'
Or handle the exit code explicitly in a wrapper script:
/usr/bin/flock -E 75 -n /var/lock/sync.lock /usr/local/bin/sync.sh
rc=$?
case "$rc" in
0) # Real Success
exit 0 ;;
75) # Intentional Lock Contention Skip (Normal)
exit 0 ;;
*) # Real Job Failure
echo "Job failed with exit code $rc" >&2
exit "$rc" ;;
esac
Pattern 4: Embedded Shell Script Locking (Self-Locking Scripts)
Instead of relying on every developer or sysadmin to remember flock in their crontab, bake the locking mechanism directly into the top of your Bash script using dedicated file descriptors:
#!/usr/bin/env bash
set -Eeuo pipefail
# ---------------------------------------------------------
# Self-Locking Boilerplate using File Descriptor 200
# ---------------------------------------------------------
LOCKFILE="/var/lock/production-backup.lock"
exec 200>"$LOCKFILE"
# Attempt non-blocking exclusive lock
if ! flock -n 200; then
echo "[$(date -u)] Another instance of $(basename "$0") is already running. Exiting cleanly." >&2
exit 0
fi
# Optional explicit unlock on script exit (Kernel handles this automatically)
trap 'flock -u 200 || true' EXIT
# ---------------------------------------------------------
# Critical Section (Guaranteed Single-Instance Execution)
# ---------------------------------------------------------
echo "[$(date -u)] Starting database backup..."
/usr/local/bin/dump-database.sh --all
echo "[$(date -u)] Backup completed successfully."
Why exec 200>"$LOCKFILE" is Superior:
exec 200>"$LOCKFILE"opens the lockfile and assigns it to File Descriptor 200 for the lifetime of the shell process.flock -n 200requests the lock directly on FD 200.- The script executes naturally. When the script exits or crashes, FD 200 closes and the kernel releases the lock automatically.
3. The 4 Fatal Flaws of Homemade Lockfiles
Homemade Lockfile Failure Modes
[ 1. TOCTOU Race ] [ 2. Stale PID Files ] [ 3. PID Recycling ] [ 4. NFS Lock Desync ]
test -f and touch are Process crashes; PID file Kernel reuses PID. Network partition splits
not atomic. Both runs remains forever. All future kill -0 returns true on lock state between hosts.
execute simultaneously. executions are blocked. completely wrong process. Both servers write data.
Flaw 1: The TOCTOU (Time-Of-Check to Time-Of-Use) Race Condition
# BROKEN: Never do this in production!
if [ ! -f /tmp/backup.lock ]; then
touch /tmp/backup.lock
/usr/local/bin/backup.sh
rm -f /tmp/backup.lock
fi
test -f and touch are two distinct, non-atomic system calls. Under concurrency spikes, two cron instances execute test -f at the exact same millisecond, observe that the file does not exist, both execute touch, and both run the critical section simultaneously.
Flaw 2: The Stale PID File Trap
# BROKEN: Leaves orphaned lockfiles on crash
echo "$$" > /var/run/app.pid
/usr/local/bin/app.sh
rm -f /var/run/app.pid
If the script encounters a segmentation fault, kernel OOM kill, or power loss, rm -f never executes. The file /var/run/app.pid sits on disk indefinitely, permanently breaking all future scheduled runs until an engineer manually deletes the file.
Flaw 3: PID Recycling False Positives
Checking kill -0 $(cat /var/run/app.pid) to verify if the process is alive fails because Linux recycles process IDs (/proc/sys/kernel/pid_max). A completely unrelated process (e.g. nginx or sshd) may inherit the old PID, causing your script to believe its predecessor is still running.
Flaw 4: Network Filesystems (NFSv3 / EFS) Incompatibilities
Local flock(2) locks are managed by the local machine's kernel. If you place a lockfile on an NFSv3 or shared network mount, lock state is not reliably propagated across multiple physical servers. Multi-server clusters require distributed locking (e.g. Redis Redlock, Postgres advisory locks pg_try_advisory_lock(), or etcd).
4. Live Diagnostics & Triage Protocol
When scheduled cron jobs appear stuck or skipped, execute these diagnostic commands:
1. View All Active Kernel Locks with lslocks
sudo lslocks | grep -Ei 'flock|backup'
- Shows the
COMMAND,PID,TYPE(FLOCK),MODE(WRITE), andPATHof the lock.
2. Find Which Process Holds the Lock with fuser and lsof
# Identify the PID holding the lockfile
sudo fuser /var/lock/production-backup.lock
# Detailed process breakdown
sudo lsof /var/lock/production-backup.lock
3. Trace File Descriptor State in /proc
# Inspect the command line of the lock holder
cat /proc/<PID>/cmdline | tr '\0' ' '
# Verify open file descriptors for FD 200
sudo ls -la /proc/<PID>/fd/ | grep -i lock
4. Detect Execution-Time Drift
If a job is constantly skipped by flock, track how long it spends in the critical section:
start=$(date +%s)
exec 200>/var/lock/job.lock
if ! flock -n 200; then
echo "[$(date -u)] Lock contention: Task taking longer than cron interval." >&2
exit 0
fi
# Run workload
/usr/local/bin/myjob.sh
duration=$(( $(date +%s) - start ))
echo "Job finished in ${duration}s"
5. How to Monitor and Alert on Skipped / Blocked Cron Jobs
While flock prevents system crashes from overlapping jobs, it introduces a subtle operational risk: The Silent Skip Problem.
If a 5-minute cron job begins taking 25 minutes to run, flock will successfully skip 4 scheduled runs. Your server won't crash, but your downstream data pipelines and backups are falling 20 minutes behind without anyone noticing.
Heartbeat Architecture for flock Tasks
[ Cron Schedule (Every 5m) ] ──► [ Acquire flock ]
│
┌───────────────────────┴───────────────────────┐
▼ ▼
Lock Contention Lock Acquired
(Deliberate Skip) │
│ ▼
▼ [ Run Workload Script ]
Exit Cleanly (0) │
(No ping sent) ┌───────────────┴───────────────┐
▼ ▼
Job Failed Job Succeeded
(Exit != 0) (Exit == 0)
│ │
▼ ▼
No Heartbeat [ Send Pingzo Ping ]
│ │
▼ ▼
[ Instant SRE Alert! ] [ Reset Health Timer ]
(WhatsApp/Slack/Telegram)
Implementing Dead Man's Switch Heartbeat Monitoring:
Use an external monitoring platform like Pingzo to track the freshness of successful executions, rather than merely tracking crontab triggers.
#!/usr/bin/env bash
set -Eeuo pipefail
LOCKFILE="/var/lock/data-pipeline.lock"
exec 200>"$LOCKFILE"
# If locked, exit immediately without sending a false success heartbeat
if ! flock -n 200; then
exit 0
fi
# Execute data pipeline
/usr/local/bin/run-data-pipeline.sh
# Send heartbeat ping ONLY if script exits with code 0 (Success)
curl -fsS -m 10 --retry 3 https://www.pingzoapp.com/api/ping/YOUR_UNIQUE_PING_SECRET
Why This Architecture Protects Production:
- No False Positives on Normal Skips: If a job is skipped because another worker is actively processing a backlog, no ping is sent.
- Dead Man's Switch Trigger: If the lock remains stuck indefinitely, or if the server crashes, no ping arrives within the expected grace window (e.g. Every 5 minutes + 2 minute grace).
- Instant Escalation: Pingzo immediately routes an alert to your on-call team via WhatsApp, Telegram, Slack, and Discord.
6. Frequently Asked Questions (FAQ)
What is the difference between flock(2) and fcntl(2) record locking?
flock(2) is designed for whole-file advisory locking and is the standard for shell scripts and cron jobs. fcntl(2) provides POSIX record/byte-range locking, allowing different processes to lock specific byte offsets within a single file (used extensively by database engines like SQLite and MySQL).
Does flock release file locks if the script crashes or gets killed by OOM?
Yes. Because flock locks are associated with the process’s open file table in the Linux kernel, the kernel automatically closes all open file descriptors and frees the lock the instant the process terminates—even on SIGKILL (kill -9) or kernel Out-Of-Memory (OOM) events.
What happens if /var/lock or /tmp is mounted on tmpfs (RAM)?
flock works perfectly on tmpfs (memory-backed filesystems). If the server reboots, the lockfile path is wiped from RAM. Because the kernel lock state is also destroyed on reboot, this poses zero stale-lock issues.
How do I run a cron job with flock without generating cron email on lock contention?
Use the -E 0 parameter with non-blocking mode: flock -E 0 -n /var/lock/app.lock -c '/path/to/script.sh'. This instructs flock to exit with status code 0 (clean exit) when lock contention occurs, preventing the cron daemon from dispatching an error email.
Is flock safe to use across multiple Docker containers sharing a bind mount?
Yes, provided both containers run on the same physical Linux host and share the same host directory via a volume bind mount. Because both containers interact with the same underlying Linux kernel VFS inode, flock will coordinate them properly. It will not work if containers are distributed across multiple separate VM nodes.
How do I handle file locks across distributed multi-server clusters?
For multi-server infrastructure, local flock is insufficient. You should use a distributed consensus lock such as Redis Redlock, PostgreSQL Advisory Locks (SELECT pg_try_advisory_lock(id)), etcd, or cloud-native schedulers like Kubernetes CronJobs with concurrencyPolicy: Forbid.
⚡ Protect Your Mission-Critical Cron Jobs Today
Prevent overlapping executions, silent job hangs, and memory exhaustion. Monitor your scheduled Linux batch tasks with Pingzo Heartbeat Monitoring and receive instant outage alerts across WhatsApp, Telegram, Slack, and Discord the second a job misses its scheduled run.
🚀 Start Monitoring Cron Jobs with Pingzo →
🛠️ Test & Generate Crontab Expressions with Free Visual Cron Builder →
Stop Finding Out About Outages from Angry Users
Get instant WhatsApp & Discord alerts the second your API, website, or server goes down. Setup in 30 seconds with 60-second checks.