Back to blog
Linux & DevOps October 2, 2026

How to Troubleshoot 502 Bad Gateway During High-Traffic Spikes

Automate WhatsApp Alerts
Start Free ➔

A 502 Bad Gateway during a flash sale, viral product launch, bot surge, or marketing campaign is rarely a simple "Nginx configuration mistake."

In reverse-proxy architectures, Nginx reports the failure because it is the boundary layer that first encounters an upstream TCP connection refusal, reset (RST), invalid FastCGI/HTTP packet, or timeout.

The underlying saturation point is usually deeper in the stack:

  • PHP-FPM worker pool exhaustion: pm.max_children saturated with requests queuing indefinitely.
  • Node.js event-loop lag: Event loop blocked on synchronous JSON parsing or CPU-heavy encryption while TCP sockets stay open.
  • Linux kernel socket queue overflows: The listen() backlog (somaxconn) or SYN backlog drops incoming handshakes.
  • Connection churn & port exhaustion: Missing upstream HTTP keepalives generating massive TIME_WAIT socket populations and ephemeral port exhaustion.
  • Database connection pool starvation: All upstream application workers stalled waiting for free database handles.
# Immediate live triage: Check socket table state and listen backlogs
ss -s
ss -lntp

30-Second Triage Table: 502 Root Causes & Emergency Actions

Symptom / Log SignatureProbable Root CauseVerification CommandPrimary Action
connect() failed (111: Connection refused)Upstream crashed, OOM killed, or listen backlog fullss -lntp, systemctl status php-fpmRestart crashed runtime, scale worker pool, raise somaxconn
recv() failed (104: Connection reset by peer)Worker thread segfault, OOM kill, or keepalive timeout racedmesg -T | grep -i oom, journalctl -u php-fpmFix memory leaks, align Nginx & Node upstream keepalive timeouts
upstream prematurely closed connectionProcess died mid-response or executed fatal execution limitCheck application fatal error logsProfile slow endpoints, inspect memory limits (memory_limit)
PHP-FPM max_children reachedConcurrency exceeds configured worker pool sizecurl http://127.0.0.1/fpm-statusIncrease pm.max_children within physical RAM limits
Node.js High Latency + Low CPUEvent loop stalled on synchronous tasks or DB pool waitingperf_hooks.monitorEventLoopDelay()Remove synchronous calls (fs.readFileSync), enlarge DB pool
too many open filesProcess or system-level File Descriptor ceiling reachedulimit -n, cat /proc/sys/fs/file-nrIncrease worker_rlimit_nofile in Nginx and systemd LimitNOFILE
Cannot assign requested addressEphemeral port range exhausted by TCP churncat /proc/sys/net/ipv4/ip_local_port_rangeEnable upstream keepalive in Nginx, expand port range
SYN_RECV count explodingSYN backlog queue full due to flash connection burstss -ant state syn-recv | wc -lIncrease net.ipv4.tcp_max_syn_backlog and net.core.somaxconn

1. The End-to-End Packet & Queue Architecture

To resolve 502 spikes during high-load events, engineers must visualize the entire request queue topology:

Client (Browser / Mobile App)
       │
       ▼ [ Edge Network / Cloudflare / WAF / ALB ]
       │
       ▼ [ Public Ingress :443 ]
┌────────────────────────────────────────────────────────┐
│ NGINX (Reverse Proxy)                                  │
│ - worker_processes auto;                               │
│ - worker_connections 16384;                            │
│ - worker_rlimit_nofile 200000;                         │
└──────────────┬─────────────────────────────────────────┘
               │
               ▼ connect() over UNIX socket or 127.0.0.1:9000
┌────────────────────────────────────────────────────────┐
│ Linux Kernel Network Subsystem                         │
│ ├── SYN Backlog (net.ipv4.tcp_max_syn_backlog)         │
│ └── Accept Queue / Listen Backlog (net.core.somaxconn) │
└──────────────┬─────────────────────────────────────────┘
               │
               ▼ accept() by Upstream Runtime
┌────────────────────────────────────────────────────────┐
│ Application Runtime Worker Pools                       │
│ ├── PHP-FPM: pm.max_children workers                   │
│ ├── Node.js: Event Loop + Worker Threads               │
│ └── Go / Python / Ruby: Puma, Gunicorn, Uvicorn        │
└──────────────┬─────────────────────────────────────────┘
               │
               ▼ Downstream Database & Microservice Queries
┌────────────────────────────────────────────────────────┐
│ Shared Dependencies                                    │
│ ├── PostgreSQL / MySQL (Connection Pool Limits)        │
│ ├── Redis Cache (Eviction / Saturation)                │
│ └── Third-Party Payment / Shipping APIs                │
└────────────────────────────────────────────────────────┘

Where Failures Manifest

  1. At the Kernel Ingress: If the SYN backlog or somaxconn accept queue fills up, incoming TCP SYN packets are silently dropped or rejected with RST, triggering client-side timeouts or Nginx 111: Connection refused.
  2. At the Worker Pool Gate: If all 80 PHP-FPM workers are processing 3-second database queries, the 81st request sits in the FPM socket listen queue. Once listen.backlog fills, Nginx aborts the connection.
  3. Inside Single-Threaded Runtimes: If a Node.js process executes JSON.parse() on a 50MB payload, the event loop pauses for 400ms. All concurrent HTTP requests in flight stall, triggering Nginx proxy_read_timeout (504) or resets (502).

2. Emergency Evidence Preservation: What to Run First

During a live outage, do not immediately restart Nginx and PHP-FPM. A brute-force restart temporarily dumps active queues, destroying forensic data regarding socket saturation, memory spikes, and slow queries.

Run this 10-second telemetry snapshot before modifying configurations:

# 1. System & Memory Pressure
uptime
free -m
vmstat 1 5

# 2. Socket Table & TCP State Summary
ss -s
ss -lnt
ss -ant state syn-recv | wc -l
ss -ant state time-wait | wc -l

# 3. Kernel Drop & Flooding Logs
dmesg -T | grep -Ei 'SYN flooding|Possible SYN flooding|oom|killed process|TCP|nf_conntrack' | tail -n 30

# 4. Nginx Error Signatures
tail -n 100 /var/log/nginx/error.log

# 5. File Descriptor Ceiling Check
cat /proc/sys/fs/file-nr
cat /proc/$(pgrep -o nginx)/limits | grep "open files"

3. Linux Kernel & Socket Layer Tuning for High Concurrency

When traffic jumps from 500 req/s to 10,000 req/s, default Linux kernel network buffers become saturated within seconds.

net.core.somaxconn and tcp_max_syn_backlog

net.core.somaxconn defines the maximum queue length for sockets in the LISTEN state awaiting accept().

# Inspect current kernel queue limits
sysctl net.core.somaxconn
sysctl net.ipv4.tcp_max_syn_backlog
sysctl net.core.netdev_max_backlog

Apply hardened production sysctl values for high-throughput edge proxies:

# /etc/sysctl.d/99-pingzo-network.conf
# Maximum listen queue backlog for socket accept()
net.core.somaxconn = 65535

# Maximum number of remembered half-open connections in SYN_RECV state
net.ipv4.tcp_max_syn_backlog = 65535

# Maximum packets queued on the network interface input queue
net.core.netdev_max_backlog = 16384

# Expand ephemeral port range to prevent local port exhaustion
net.ipv4.ip_local_port_range = 10240 65535

# Enable TCP window scaling and fast recycle of FIN-WAIT sockets
net.ipv4.tcp_fin_timeout = 15
# Reload sysctl without rebooting
sudo sysctl --system

[!CAUTION] Avoid setting net.ipv4.tcp_tw_reuse = 1 blindly on modern Linux kernels unless you have verified routing semantics. In Linux kernels $\ge$ 5.x, tcp_tw_reuse = 2 is enabled for loopback traffic by default. Misconfiguring this setting across stateful NAT firewalls can lead to dropped SYN packets and corrupted TCP streams.


4. Nginx Tuning: Upstream Connection Pooling & Keepalives

The most prevalent architectural bug in Nginx reverse proxies during traffic spikes is upstream connection churn.

By default, Nginx creates a brand-new TCP connection to the upstream application (PHP-FPM, Node.js, Go) for every single incoming HTTP request, tearing it down with a FIN/RST sequence immediately after the response. Under 5,000 req/s, this triggers massive CPU overhead and floods the host with tens of thousands of TIME_WAIT sockets.

Enabling HTTP/1.1 Upstream Keepalives

Configure persistent connection pooling in your upstream blocks:

# /etc/nginx/conf.d/upstream_app.conf
upstream nodejs_backend {
    server 127.0.0.1:3000 max_fails=3 fail_timeout=10s;
    
    # Cache up to 128 idle keepalive connections per Nginx worker process
    keepalive 128;
    keepalive_requests 10000;
    keepalive_timeout 60s;
}

server {
    listen 443 ssl http2;
    server_name api.example.com;

    location / {
        # CRITICAL: Force HTTP/1.1 and clear Connection header for keepalive reuse
        proxy_http_version 1.1;
        proxy_set_header Connection "";
        
        proxy_set_header Host $host;
        proxy_set_header X-Real-IP $remote_addr;
        proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
        proxy_set_header X-Forwarded-Proto $scheme;

        # Fine-tuned upstream timeouts
        proxy_connect_timeout 3s;
        proxy_send_timeout 30s;
        proxy_read_timeout 30s;

        # Buffer sizing to prevent disk spools
        proxy_buffer_size 16k;
        proxy_buffers 8 16k;
        proxy_busy_buffers_size 32k;

        proxy_pass http://nodejs_backend;
    }
}

Validate your live response headers and proxy routing with the Pingzo HTTP Header Checker.


5. PHP-FPM Sizing & Capacity Engineering

PHP-FPM is a process-based runtime where each worker process consumes a fixed slice of server RAM (typically 60MB to 200MB per worker depending on framework dependencies).

The Worker Sizing Formula

Never guess pm.max_children. Use real memory measurements:

$$\text{pm.max_children} = \frac{\text{Total Available RAM} - \text{System Reserve (OS + Nginx + DB)}}{\text{Average PHP Worker RSS Memory}}$$

# Calculate average RSS memory consumed per PHP-FPM worker (in MB)
ps --no-headers -o rss,cmd -C php-fpm | awk '{ sum+=$1; count++ } END { if (count>0) print "Avg Worker RSS: " sum/count/1024 " MB (Workers: " count ")" }'

If your server has 16 GB RAM, reserving 4 GB for the OS and Nginx leaves 12 GB ($12,288\text{ MB}$) for PHP. If average worker RSS is 120 MB:

$$\text{pm.max_children} = \frac{12288\text{ MB}}{120\text{ MB}} \approx 102\text{ workers}$$

Hardened php-fpm.conf Pool Configuration

; /etc/php/8.3/fpm/pool.d/www.conf
[www]
user = www-data
group = www-data

; Use UNIX sockets for local Nginx-to-FPM communication (reduces TCP overhead)
listen = /run/php/php8.3-fpm.sock
listen.owner = www-data
listen.group = www-data
listen.mode = 0660

; Align socket listen backlog with kernel somaxconn
listen.backlog = 65535

; Use static process management for consistent high-throughput servers
; (Eliminates runtime fork() latency during sudden traffic spikes)
pm = dynamic
pm.max_children = 100
pm.start_servers = 20
pm.min_spare_servers = 10
pm.max_spare_servers = 40
pm.max_requests = 1000

; Enable slow logging to pinpoint blocking database queries during spikes
request_slowlog_timeout = 3s
slowlog = /var/log/php-fpm/www-slow.log

; Enable the internal status page for real-time telemetry
pm.status_path = /fpm-status

Inspecting PHP-FPM Status During Load

Query the FPM status endpoint locally:

curl -s http://127.0.0.1/fpm-status?json | jq .

Key saturation indicators:

  • active processes: If this equals pm.max_children, your worker pool is 100% saturated.
  • listen queue: If $> 0$, incoming requests are queuing in the socket buffer.
  • max children reached: Counter increments every time a request is delayed due to worker starvation.

6. Node.js Event-Loop Starvation & Concurrency Tuning

Unlike PHP-FPM, Node.js runs on a single-threaded JavaScript event loop. A Node.js instance can return 502/504 errors even while total server CPU utilization sits at only 25% (on an 8-core machine where one core is 100% blocked).

Measuring Event-Loop Delay

Inject native event loop telemetry using node:perf_hooks:

// server.js
import http from 'node:http';
import { monitorEventLoopDelay } from 'node:perf_hooks';

// Resolution: sample every 20ms
const histogram = monitorEventLoopDelay({ resolution: 20 });
histogram.enable();

setInterval(() => {
  const p99_ms = histogram.percentile(99) / 1e6;
  const max_ms = histogram.max / 1e6;
  
  if (p99_ms > 100) {
    console.warn(`[EVENT LOOP LAG] p99=${p99_ms.toFixed(2)}ms max=${max_ms.toFixed(2)}ms`);
  }
  histogram.reset();
}, 5000);

Eliminating Synchronous Blockers

  1. Replace Synchronous I/O: Never use fs.readFileSync(), crypto.pbkdf2Sync(), or large synchronous regex evaluations in active request routes.
  2. Cluster Across CPU Cores: Use Node.js cluster or a process manager like PM2:
# Run Node.js clustered across all available CPU cores
pm2 start ecosystem.config.js -i max
  1. Align Keepalive Timeouts: Ensure Node's server.keepAliveTimeout is longer than Nginx's proxy_read_timeout to prevent race conditions where Node closes a socket just as Nginx reuses it.
const server = app.listen(3000);
server.keepAliveTimeout = 65000; // 65 seconds (Nginx is 60s)
server.headersTimeout = 66000;

7. Granular Latency Profiling: Pinpointing the Delay Domain

When a 502 occurs, determine whether the stall happens during DNS, TCP handshake, or upstream execution.

Run curl with detailed format metrics:

curl -sS -o /dev/null -w "
HTTP Code:          %{http_code}
DNS Lookup:         %{time_namelookup}s
TCP Handshake:      %{time_connect}s
TLS Handshake:      %{time_appconnect}s
Time to First Byte: %{time_starttransfer}s
Total Time:         %{time_total}s
" https://api.example.com/health

Metric Interpretation:

  • time_connect High ($> 1.0\text{s}$): Network congestion, firewall rate limits, or kernel somaxconn / SYN backlog dropping packets.
  • time_starttransfer High ($> 5.0\text{s}$): Network and TLS succeeded immediately, but the application runtime (PHP/Node) or database query took seconds to compute the first byte.

Test HTTP status codes and response headers across edge networks using the Pingzo HTTP Status Code Checker.


8. Automated Load Testing with k6

Before running major traffic campaigns, simulate concurrency limits in staging to discover the exact threshold where 502 errors begin.

// load-test.js
import http from 'k6/http';
import { check, sleep } from 'k6';

export const options = {
  stages: [
    { duration: '1m', target: 200 },   // Warm up to 200 virtual users
    { duration: '2m', target: 1000 },  // Spike to 1,000 users
    { duration: '3m', target: 3000 },  // Push to peak 3,000 users
    { duration: '1m', target: 0 },     // Ramp down
  ],
  thresholds: {
    http_req_failed: ['rate<0.01'],    // Error rate must be under 1%
    http_req_duration: ['p(95)<500'],  // 95% of requests must complete in under 500ms
  },
};

export default function () {
  const res = http.get('https://staging.example.com/api/products');
  check(res, {
    'status is 200': (r) => r.status === 200,
    'not 502/504': (r) => r.status !== 502 && r.status !== 504,
  });
  sleep(0.5);
}
# Execute load test
k6 run load-test.js

9. Synthetic Edge Health & Upstream Reset Monitoring

Internal server-side CPU metrics often report normal health while external users experience severe 502 outages because the failure happens at edge ingress or reverse proxy handoffs.

  1. Global Multi-Region Probing: Deploy external synthetic probes via Pingzo Uptime Monitoring to poll your core APIs every 30 seconds from multiple geographic regions.
  2. Dedicated 502/504 Threshold Alerts: Configure P1 alerts specifically when the ratio of 502 Bad Gateway responses exceeds 0.5% of total edge traffic over a 2-minute rolling window.
  3. Database Pool Saturation Early Warning: Alert on database connection pool utilization exceeding 80% before worker starvation cascades into Nginx gateway timeouts.

Monitor your server reachability and packet latency using the free Pingzo Ping Test.


Frequently Asked Questions

Why did increasing proxy_read_timeout not fix my 502?

proxy_read_timeout only sets how long Nginx will wait for an upstream process to send a response block. If your PHP-FPM workers or Node event loops are 100% saturated, increasing the timeout merely keeps stalled connections open longer, consuming more file descriptors and memory until the server crashes.

What is the difference between a 502 Bad Gateway and a 504 Gateway Timeout?

A 502 Bad Gateway means Nginx contacted the upstream, but the upstream actively rejected the connection (Connection refused), crashed mid-stream (Connection reset by peer), or returned an invalid header payload. A 504 Gateway Timeout means Nginx successfully connected to the upstream, but the upstream failed to send data within the configured proxy_read_timeout limit.

Why do 502 errors disappear immediately when restarting PHP-FPM?

Restarting PHP-FPM forcefully terminates all existing stalled worker processes and flushes the socket listen backlog. Active requests succeed immediately on the fresh worker pool. However, if the underlying cause (slow database queries, unindexed tables, insufficient pm.max_children) is not addressed, workers will saturate again within minutes as traffic resumes.

Zero-Code Uptime Alerts

Stop Finding Out About Outages from Angry Users

Get instant WhatsApp & Discord alerts the second your API, website, or server goes down. Setup in 30 seconds with 60-second checks.

WhatsApp & Discord 60-Second Checks Free Forever Plan
Try Pingzo Free

Know before your users do

Connect official WhatsApp notification channels, Discord webhooks, Telegram bots, and public status pages. Start in 30 seconds.

Create Free Monitor