The independent field guide for on-call teams
Build an on-call system people can trust.
Practical guides, assessments, and templates for alert quality, escalation, fair rotations, handoffs, compensation, and incident response.
Free tools. No signup required. Written for the people who carry the pager.
Detected
API latency high
Routed
Payments primary
No acknowledgement
5 min
Escalated
Platform secondary
Owned by a human
Acknowledged
Before the incident
On-call fails quietly before it fails loudly.
Coverage gaps, noisy alerts, unclear ownership, and brittle escalation paths usually exist long before the night they become urgent. A good on-call system makes those risks visible while there is still time to fix them.
Coverage gaps
Know who owns each hour, including holidays, leave, and handoffs.
Noisy alerts
Separate signals that require action from notifications that merely create interruption.
Ambiguous escalation
Define what happens when the first responder cannot acknowledge or resolve the issue.
Uneven load
Measure nights, weekends, interruptions, and recovery—not only total shift hours.
Learn it. Plan it. Improve it.
Learn
Clear guides for the decisions behind schedules, escalation, alerting, handoffs, and incident response.
Browse guidesPlan
Score alert fatigue, draft an escalation path, assess readiness, and export practical team artifacts.
Open the toolsImprove
Score your current setup, identify the biggest risks, and turn them into a focused 30-day plan.
Check your readinessUseful before the pager rings.
On-Call Readiness Assessment
Score coverage, escalation, alert quality, documentation, fairness, testing, and notification delivery.
Get your scoreAlert Fatigue Score
Estimate how much paging load is actionable, duplicated, auto-resolved, or missing context.
Check alert noiseEscalation Policy Builder
Turn a verbal fallback plan into explicit steps, timeouts, targets, and stop conditions.
Build a policyOn-Call Policy Template
Make coverage, handoffs, compensation, recovery, escalation, and ownership explicit before the pager rings.
Use the template
A practical path from signal to response.
01
Detect
A monitoring system identifies a condition that may need action.
02
Route
Ownership rules choose the responsible service, team, and path.
03
Page
The current responder receives the configured notification.
04
Acknowledge
The system records that a human has taken ownership.
05
Escalate or resolve
The path continues until the incident has a responsible next step.
06
Learn
The team improves alerts, runbooks, or ownership based on what happened.
Reliability includes the responder.
A schedule can be technically complete and still be unfair, exhausting, or impossible to sustain. Strong on-call programs make room for coverage requests, shadowing, recovery after disrupted nights, clear compensation, and escalation without blame.
Moving providers
Migrate the operating model, not only the configuration.
A safe migration includes schedules, escalation rules, contact methods, integrations, ownership, runbooks, audit requirements, and a tested cutover path.
Opsgenie migration checklist
Inventory, parity matrix, dual-run plan, controlled cutover test, and rollback.
Grafana OnCall OSS alternatives
What changed, what still works locally, and how to choose a replacement.
Provider evaluation worksheet
Turn requirements into an explicit, testable escalation path before you compare vendors.
FROM THE TEAM BEHIND ONCALL.FYI
Need a simpler path from critical signal to notification?
MonoDuty brings webhook alerting, uptime monitoring, and check-in/heartbeat alerts into one focused workflow. Use it when the gap is reliable detection and notification delivery—not when you need a full incident-management suite.
MonoDuty and oncall.fyi are built by the same team.
The On-Call Brief
One practical idea for calmer, fairer incident response.
Short field notes, templates, and operational lessons. No daily noise and no incident theater.