Back to Learn Hub
On-Call & Alerting July 23, 2026

The On-Call Onboarding 30-Day Checklist for SREs | Pingzo

SSumit Nath

The On-Call Onboarding 30-Day Checklist

Joining an active on-call rotation can be one of the most stressful experiences for a new developer or Site Reliability Engineer (SRE). Being woken up at 3:00 AM to resolve a production system issue requires deep familiarity with the stack, alerting rules, and runbooks.

To prevent burnout and ensure a smooth transition, teams should use a structured onboarding checklist. This 30-day playbook helps new engineers build confidence and learn operational procedures before taking solo shifts.


📅 Week 1: Alert Configurations & Tooling Setup

Focus on configuring communication tools and establishing critical notification overrides:

  • Configure Alert Channel Overrides: Set up your phone to bypass silent mode (Do Not Disturb) for critical alerts. On iOS/Android, configure emergency contacts or use Pingzo's official WhatsApp alert channel.
  • Verify Authentication Access: Ensure you have verified admin access to target servers, databases, third-party status pages, and monitoring consoles.
  • Locate the Service Catalog: Review the list of active backend services, databases, and cron tasks.
  • Join Communication Channels: Request access to SRE discussion rooms (Slack, Discord, or Teams) where alerts are actively triaged.

📅 Week 2: Shadowing & Runbook Review

Begin shadowing active on-call engineers to observe real-world incident management:

  • Shadow the Primary On-Call: Join active incident triage calls. Observe how alerts are acknowledged, status updates are published, and diagnostic commands are run.
  • Review Recent Incident Reports: Read the post-mortem documents from the last 3 to 6 months to learn about common architectural bottlenecks.
  • Review Escalation Policies: Learn when to escalate a database failover, an API slowdown, or an SSL expiration warning to senior team members.
  • Verify Runbook Access: Ensure every active HTTP and cron monitor in your dashboard links to a valid runbook containing recovery instructions.

📅 Week 3: Supervised Secondary On-Call

Take on active alert triage under the direct supervision of an experienced shadow buddy:

  • Act as the Secondary On-Call: Receive alerts alongside the primary engineer. Perform initial analysis and propose mitigation actions.
  • Acknowledge and Resolve Test Incidents: Run a scheduled failover test to practice acknowledging and resolving alerts.
  • Draft Incident Resolutions: Practice writing clear resolution notes in Pingzo when closing resolved outages.

📅 Week 4: The Solo Shift

Safely take on your first primary on-call rotation:

  • Begin Solo On-Call Rotation: Serve as the first line of defense for system alerts.
  • Follow Runbooks to Resolve Alerts: Use pre-defined recovery scripts and instructions linked in Pingzo monitors.
  • Conduct Blameless Reviews: Lead the post-incident review meeting for any outages that occurred during your shift.