Monitoring tools can detect outages instantly. However, if those alerts are routed to the wrong developers, or if every minor warning pages the entire engineering team, you will suffer from alert fatigue.
Responder burnout leads to delayed response times and missed critical outages. To protect developer schedules and resolve incidents quickly, you must establish an Incident Escalation Policy.
This guide details best practices for structuring escalation policies for growing teams.
🛠️ The Three-Tier Escalation Model
A standard incident escalation policy routes pages sequentially based on responder availability:
graph TD
A[Outage Detected] --> B[Tier 1: Primary On-Call]
B -- Acknowledged within 5m? --> C{Yes / No}
C -- Yes --> D[Resolution Triage]
C -- No --> E[Tier 2: Secondary On-Call]
E -- Acknowledged within 10m? --> F{Yes / No}
F -- Yes --> D
F -- No --> G[Tier 3: Engineering Director / Stakeholders]
Tier 1: The Primary On-Call Engineer
The primary responder is the first line of defense. They receive all high-severity pages during their shift. If they acknowledge the alert within 5 minutes, the escalation chain pauses.
Tier 2: The Secondary On-Call (The Shadow)
If the primary responder fails to acknowledge the page within the target time (e.g., they are asleep or out of cell range), the alert escalates to the secondary engineer.
Tier 3: Management / Stakeholder Escalation
If both the primary and secondary responders fail to acknowledge the incident within 15 minutes, the alert escalates to the engineering director or VP to ensure critical customer outages are triaged.
🚀 Escalation Policy Best Practices
1. Separate Notifications by Severity
Never route non-critical warnings to high-priority alert channels. P3 and P4 issues should go to standard Slack channels, reserving SMS, phone calls, and WhatsApp notifications for P1 and P2 outages that require immediate response.
2. Automate Roster Rotations
Use scheduling tools to rotate primary and secondary duties automatically every week, preventing on-call responsibilities from falling on a small group of developers.
3. Maintain Blameless Escalation Reviews
If an alert frequently escalates to Tier 2, review the incident in your post-mortem. Was the primary engineer experiencing alert fatigue? Were alert volume thresholds set too low? Focus on adjusting configurations rather than assigning blame.