A 502 Bad Gateway during a flash sale, viral product launch, bot surge, or marketing campaign is rarely a simple "Nginx configuration mistake."
In reverse-proxy architectures, Nginx reports the failure because it is the boundary layer that first encounters an upstream TCP connection refusal, reset (RST), invalid FastCGI/HTTP packet, or timeout.
The underlying saturation point is usually deeper in the stack:
- PHP-FPM worker pool exhaustion:
pm.max_childrensaturated with requests queuing indefinitely. - Node.js event-loop lag: Event loop blocked on synchronous JSON parsing or CPU-heavy encryption while TCP sockets stay open.
- Linux kernel socket queue overflows: The
listen()backlog (somaxconn) or SYN backlog drops incoming handshakes. - Connection churn & port exhaustion: Missing upstream HTTP keepalives generating massive
TIME_WAITsocket populations and ephemeral port exhaustion. - Database connection pool starvation: All upstream application workers stalled waiting for free database handles.
# Immediate live triage: Check socket table state and listen backlogs
ss -s
ss -lntp
30-Second Triage Table: 502 Root Causes & Emergency Actions
| Symptom / Log Signature | Probable Root Cause | Verification Command | Primary Action |
|---|---|---|---|
connect() failed (111: Connection refused) | Upstream crashed, OOM killed, or listen backlog full | ss -lntp, systemctl status php-fpm | Restart crashed runtime, scale worker pool, raise somaxconn |
recv() failed (104: Connection reset by peer) | Worker thread segfault, OOM kill, or keepalive timeout race | dmesg -T | grep -i oom, journalctl -u php-fpm | Fix memory leaks, align Nginx & Node upstream keepalive timeouts |
upstream prematurely closed connection | Process died mid-response or executed fatal execution limit | Check application fatal error logs | Profile slow endpoints, inspect memory limits (memory_limit) |
PHP-FPM max_children reached | Concurrency exceeds configured worker pool size | curl http://127.0.0.1/fpm-status | Increase pm.max_children within physical RAM limits |
| Node.js High Latency + Low CPU | Event loop stalled on synchronous tasks or DB pool waiting | perf_hooks.monitorEventLoopDelay() | Remove synchronous calls (fs.readFileSync), enlarge DB pool |
too many open files | Process or system-level File Descriptor ceiling reached | ulimit -n, cat /proc/sys/fs/file-nr | Increase worker_rlimit_nofile in Nginx and systemd LimitNOFILE |
Cannot assign requested address | Ephemeral port range exhausted by TCP churn | cat /proc/sys/net/ipv4/ip_local_port_range | Enable upstream keepalive in Nginx, expand port range |
SYN_RECV count exploding | SYN backlog queue full due to flash connection burst | ss -ant state syn-recv | wc -l | Increase net.ipv4.tcp_max_syn_backlog and net.core.somaxconn |
1. The End-to-End Packet & Queue Architecture
To resolve 502 spikes during high-load events, engineers must visualize the entire request queue topology:
Client (Browser / Mobile App)
│
▼ [ Edge Network / Cloudflare / WAF / ALB ]
│
▼ [ Public Ingress :443 ]
┌────────────────────────────────────────────────────────┐
│ NGINX (Reverse Proxy) │
│ - worker_processes auto; │
│ - worker_connections 16384; │
│ - worker_rlimit_nofile 200000; │
└──────────────┬─────────────────────────────────────────┘
│
▼ connect() over UNIX socket or 127.0.0.1:9000
┌────────────────────────────────────────────────────────┐
│ Linux Kernel Network Subsystem │
│ ├── SYN Backlog (net.ipv4.tcp_max_syn_backlog) │
│ └── Accept Queue / Listen Backlog (net.core.somaxconn) │
└──────────────┬─────────────────────────────────────────┘
│
▼ accept() by Upstream Runtime
┌────────────────────────────────────────────────────────┐
│ Application Runtime Worker Pools │
│ ├── PHP-FPM: pm.max_children workers │
│ ├── Node.js: Event Loop + Worker Threads │
│ └── Go / Python / Ruby: Puma, Gunicorn, Uvicorn │
└──────────────┬─────────────────────────────────────────┘
│
▼ Downstream Database & Microservice Queries
┌────────────────────────────────────────────────────────┐
│ Shared Dependencies │
│ ├── PostgreSQL / MySQL (Connection Pool Limits) │
│ ├── Redis Cache (Eviction / Saturation) │
│ └── Third-Party Payment / Shipping APIs │
└────────────────────────────────────────────────────────┘
Where Failures Manifest
- At the Kernel Ingress: If the SYN backlog or
somaxconnaccept queue fills up, incoming TCP SYN packets are silently dropped or rejected withRST, triggering client-side timeouts or Nginx111: Connection refused. - At the Worker Pool Gate: If all 80 PHP-FPM workers are processing 3-second database queries, the 81st request sits in the FPM socket listen queue. Once
listen.backlogfills, Nginx aborts the connection. - Inside Single-Threaded Runtimes: If a Node.js process executes
JSON.parse()on a 50MB payload, the event loop pauses for 400ms. All concurrent HTTP requests in flight stall, triggering Nginxproxy_read_timeout(504) or resets (502).
2. Emergency Evidence Preservation: What to Run First
During a live outage, do not immediately restart Nginx and PHP-FPM. A brute-force restart temporarily dumps active queues, destroying forensic data regarding socket saturation, memory spikes, and slow queries.
Run this 10-second telemetry snapshot before modifying configurations:
# 1. System & Memory Pressure
uptime
free -m
vmstat 1 5
# 2. Socket Table & TCP State Summary
ss -s
ss -lnt
ss -ant state syn-recv | wc -l
ss -ant state time-wait | wc -l
# 3. Kernel Drop & Flooding Logs
dmesg -T | grep -Ei 'SYN flooding|Possible SYN flooding|oom|killed process|TCP|nf_conntrack' | tail -n 30
# 4. Nginx Error Signatures
tail -n 100 /var/log/nginx/error.log
# 5. File Descriptor Ceiling Check
cat /proc/sys/fs/file-nr
cat /proc/$(pgrep -o nginx)/limits | grep "open files"
3. Linux Kernel & Socket Layer Tuning for High Concurrency
When traffic jumps from 500 req/s to 10,000 req/s, default Linux kernel network buffers become saturated within seconds.
net.core.somaxconn and tcp_max_syn_backlog
net.core.somaxconn defines the maximum queue length for sockets in the LISTEN state awaiting accept().
# Inspect current kernel queue limits
sysctl net.core.somaxconn
sysctl net.ipv4.tcp_max_syn_backlog
sysctl net.core.netdev_max_backlog
Apply hardened production sysctl values for high-throughput edge proxies:
# /etc/sysctl.d/99-pingzo-network.conf
# Maximum listen queue backlog for socket accept()
net.core.somaxconn = 65535
# Maximum number of remembered half-open connections in SYN_RECV state
net.ipv4.tcp_max_syn_backlog = 65535
# Maximum packets queued on the network interface input queue
net.core.netdev_max_backlog = 16384
# Expand ephemeral port range to prevent local port exhaustion
net.ipv4.ip_local_port_range = 10240 65535
# Enable TCP window scaling and fast recycle of FIN-WAIT sockets
net.ipv4.tcp_fin_timeout = 15
# Reload sysctl without rebooting
sudo sysctl --system
[!CAUTION] Avoid setting
net.ipv4.tcp_tw_reuse = 1blindly on modern Linux kernels unless you have verified routing semantics. In Linux kernels $\ge$ 5.x,tcp_tw_reuse = 2is enabled for loopback traffic by default. Misconfiguring this setting across stateful NAT firewalls can lead to dropped SYN packets and corrupted TCP streams.
4. Nginx Tuning: Upstream Connection Pooling & Keepalives
The most prevalent architectural bug in Nginx reverse proxies during traffic spikes is upstream connection churn.
By default, Nginx creates a brand-new TCP connection to the upstream application (PHP-FPM, Node.js, Go) for every single incoming HTTP request, tearing it down with a FIN/RST sequence immediately after the response. Under 5,000 req/s, this triggers massive CPU overhead and floods the host with tens of thousands of TIME_WAIT sockets.
Enabling HTTP/1.1 Upstream Keepalives
Configure persistent connection pooling in your upstream blocks:
# /etc/nginx/conf.d/upstream_app.conf
upstream nodejs_backend {
server 127.0.0.1:3000 max_fails=3 fail_timeout=10s;
# Cache up to 128 idle keepalive connections per Nginx worker process
keepalive 128;
keepalive_requests 10000;
keepalive_timeout 60s;
}
server {
listen 443 ssl http2;
server_name api.example.com;
location / {
# CRITICAL: Force HTTP/1.1 and clear Connection header for keepalive reuse
proxy_http_version 1.1;
proxy_set_header Connection "";
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
# Fine-tuned upstream timeouts
proxy_connect_timeout 3s;
proxy_send_timeout 30s;
proxy_read_timeout 30s;
# Buffer sizing to prevent disk spools
proxy_buffer_size 16k;
proxy_buffers 8 16k;
proxy_busy_buffers_size 32k;
proxy_pass http://nodejs_backend;
}
}
Validate your live response headers and proxy routing with the Pingzo HTTP Header Checker.
5. PHP-FPM Sizing & Capacity Engineering
PHP-FPM is a process-based runtime where each worker process consumes a fixed slice of server RAM (typically 60MB to 200MB per worker depending on framework dependencies).
The Worker Sizing Formula
Never guess pm.max_children. Use real memory measurements:
$$\text{pm.max_children} = \frac{\text{Total Available RAM} - \text{System Reserve (OS + Nginx + DB)}}{\text{Average PHP Worker RSS Memory}}$$
# Calculate average RSS memory consumed per PHP-FPM worker (in MB)
ps --no-headers -o rss,cmd -C php-fpm | awk '{ sum+=$1; count++ } END { if (count>0) print "Avg Worker RSS: " sum/count/1024 " MB (Workers: " count ")" }'
If your server has 16 GB RAM, reserving 4 GB for the OS and Nginx leaves 12 GB ($12,288\text{ MB}$) for PHP. If average worker RSS is 120 MB:
$$\text{pm.max_children} = \frac{12288\text{ MB}}{120\text{ MB}} \approx 102\text{ workers}$$
Hardened php-fpm.conf Pool Configuration
; /etc/php/8.3/fpm/pool.d/www.conf
[www]
user = www-data
group = www-data
; Use UNIX sockets for local Nginx-to-FPM communication (reduces TCP overhead)
listen = /run/php/php8.3-fpm.sock
listen.owner = www-data
listen.group = www-data
listen.mode = 0660
; Align socket listen backlog with kernel somaxconn
listen.backlog = 65535
; Use static process management for consistent high-throughput servers
; (Eliminates runtime fork() latency during sudden traffic spikes)
pm = dynamic
pm.max_children = 100
pm.start_servers = 20
pm.min_spare_servers = 10
pm.max_spare_servers = 40
pm.max_requests = 1000
; Enable slow logging to pinpoint blocking database queries during spikes
request_slowlog_timeout = 3s
slowlog = /var/log/php-fpm/www-slow.log
; Enable the internal status page for real-time telemetry
pm.status_path = /fpm-status
Inspecting PHP-FPM Status During Load
Query the FPM status endpoint locally:
curl -s http://127.0.0.1/fpm-status?json | jq .
Key saturation indicators:
active processes: If this equalspm.max_children, your worker pool is 100% saturated.listen queue: If $> 0$, incoming requests are queuing in the socket buffer.max children reached: Counter increments every time a request is delayed due to worker starvation.
6. Node.js Event-Loop Starvation & Concurrency Tuning
Unlike PHP-FPM, Node.js runs on a single-threaded JavaScript event loop. A Node.js instance can return 502/504 errors even while total server CPU utilization sits at only 25% (on an 8-core machine where one core is 100% blocked).
Measuring Event-Loop Delay
Inject native event loop telemetry using node:perf_hooks:
// server.js
import http from 'node:http';
import { monitorEventLoopDelay } from 'node:perf_hooks';
// Resolution: sample every 20ms
const histogram = monitorEventLoopDelay({ resolution: 20 });
histogram.enable();
setInterval(() => {
const p99_ms = histogram.percentile(99) / 1e6;
const max_ms = histogram.max / 1e6;
if (p99_ms > 100) {
console.warn(`[EVENT LOOP LAG] p99=${p99_ms.toFixed(2)}ms max=${max_ms.toFixed(2)}ms`);
}
histogram.reset();
}, 5000);
Eliminating Synchronous Blockers
- Replace Synchronous I/O: Never use
fs.readFileSync(),crypto.pbkdf2Sync(), or large synchronous regex evaluations in active request routes. - Cluster Across CPU Cores: Use Node.js
clusteror a process manager like PM2:
# Run Node.js clustered across all available CPU cores
pm2 start ecosystem.config.js -i max
- Align Keepalive Timeouts: Ensure Node's
server.keepAliveTimeoutis longer than Nginx'sproxy_read_timeoutto prevent race conditions where Node closes a socket just as Nginx reuses it.
const server = app.listen(3000);
server.keepAliveTimeout = 65000; // 65 seconds (Nginx is 60s)
server.headersTimeout = 66000;
7. Granular Latency Profiling: Pinpointing the Delay Domain
When a 502 occurs, determine whether the stall happens during DNS, TCP handshake, or upstream execution.
Run curl with detailed format metrics:
curl -sS -o /dev/null -w "
HTTP Code: %{http_code}
DNS Lookup: %{time_namelookup}s
TCP Handshake: %{time_connect}s
TLS Handshake: %{time_appconnect}s
Time to First Byte: %{time_starttransfer}s
Total Time: %{time_total}s
" https://api.example.com/health
Metric Interpretation:
time_connectHigh ($> 1.0\text{s}$): Network congestion, firewall rate limits, or kernelsomaxconn/ SYN backlog dropping packets.time_starttransferHigh ($> 5.0\text{s}$): Network and TLS succeeded immediately, but the application runtime (PHP/Node) or database query took seconds to compute the first byte.
Test HTTP status codes and response headers across edge networks using the Pingzo HTTP Status Code Checker.
8. Automated Load Testing with k6
Before running major traffic campaigns, simulate concurrency limits in staging to discover the exact threshold where 502 errors begin.
// load-test.js
import http from 'k6/http';
import { check, sleep } from 'k6';
export const options = {
stages: [
{ duration: '1m', target: 200 }, // Warm up to 200 virtual users
{ duration: '2m', target: 1000 }, // Spike to 1,000 users
{ duration: '3m', target: 3000 }, // Push to peak 3,000 users
{ duration: '1m', target: 0 }, // Ramp down
],
thresholds: {
http_req_failed: ['rate<0.01'], // Error rate must be under 1%
http_req_duration: ['p(95)<500'], // 95% of requests must complete in under 500ms
},
};
export default function () {
const res = http.get('https://staging.example.com/api/products');
check(res, {
'status is 200': (r) => r.status === 200,
'not 502/504': (r) => r.status !== 502 && r.status !== 504,
});
sleep(0.5);
}
# Execute load test
k6 run load-test.js
9. Synthetic Edge Health & Upstream Reset Monitoring
Internal server-side CPU metrics often report normal health while external users experience severe 502 outages because the failure happens at edge ingress or reverse proxy handoffs.
- Global Multi-Region Probing: Deploy external synthetic probes via Pingzo Uptime Monitoring to poll your core APIs every 30 seconds from multiple geographic regions.
- Dedicated 502/504 Threshold Alerts: Configure P1 alerts specifically when the ratio of
502 Bad Gatewayresponses exceeds 0.5% of total edge traffic over a 2-minute rolling window. - Database Pool Saturation Early Warning: Alert on database connection pool utilization exceeding 80% before worker starvation cascades into Nginx gateway timeouts.
Monitor your server reachability and packet latency using the free Pingzo Ping Test.
Frequently Asked Questions
Why did increasing proxy_read_timeout not fix my 502?
proxy_read_timeout only sets how long Nginx will wait for an upstream process to send a response block. If your PHP-FPM workers or Node event loops are 100% saturated, increasing the timeout merely keeps stalled connections open longer, consuming more file descriptors and memory until the server crashes.
What is the difference between a 502 Bad Gateway and a 504 Gateway Timeout?
A 502 Bad Gateway means Nginx contacted the upstream, but the upstream actively rejected the connection (Connection refused), crashed mid-stream (Connection reset by peer), or returned an invalid header payload. A 504 Gateway Timeout means Nginx successfully connected to the upstream, but the upstream failed to send data within the configured proxy_read_timeout limit.
Why do 502 errors disappear immediately when restarting PHP-FPM?
Restarting PHP-FPM forcefully terminates all existing stalled worker processes and flushes the socket listen backlog. Active requests succeed immediately on the fresh worker pool. However, if the underlying cause (slow database queries, unindexed tables, insufficient pm.max_children) is not addressed, workers will saturate again within minutes as traffic resumes.
Stop Finding Out About Outages from Angry Users
Get instant WhatsApp & Discord alerts the second your API, website, or server goes down. Setup in 30 seconds with 60-second checks.