What is an escalation policy?
An escalation policy is the set of rules that decides who gets paged for an alert, in what order, and how long each responder has to acknowledge before the next person is paged. A good policy always ends with someone who will answer, so an alert is never left without a responder.
How an escalation policy works
A policy is an ordered list of steps. Each step names who to page, how to reach them, and how long to wait for an acknowledgement. When an alert is routed to the policy, the first step pages. If nobody acknowledges within the timeout, the next step pages. When the last step is reached without an answer, the policy either repeats from the top or stops, depending on how it is set.
Steps usually point at a schedule rather than a named person, so the policy pages whoever is on call right now. A step can also page a channel, a team lead, or several people at once. The moment one person acknowledges, escalation stops and that person owns the alert.
Three things that get confused
Three parts of an on-call setup are often called by the same name. Keeping them apart makes each one easier to change.
| Part | The question it answers | Who normally owns it |
|---|---|---|
| Schedule | Who is on call, and when? | The team that carries it |
| Routing | Which policy does this alert go to? | The team that owns the service |
| Escalation policy | Who is paged, how, and what if they do not answer? | The team that owns the service |
A schedule is a calendar of shifts and handoffs. Routing rules read an alert’s service, severity or source and choose a policy. The escalation policy is the sequence that runs once an alert has been routed to it. If a change to one of the three forces a change to the others, the boundaries between them are in the wrong place.
A worked example
This is an illustrative policy, not a recommendation for every team. A payments team might run one policy for its critical alerts:
| Step | Who is paged, and how | Wait before the next step |
|---|---|---|
| 1 | The primary on-call engineer, by phone call and SMS | 5 minutes |
| 2 | The secondary on-call engineer, by phone call and SMS, and the channel | 10 minutes |
| 3 | The engineering manager on call, by phone call and SMS | 15 minutes |
| 4 | Repeat from step 1 |
Non-critical alerts from the same services go to a second policy that messages the channel during working hours and pages nobody at night. The two policies differ in urgency, not in who owns them.
What makes a good policy
- It always ends with someone who answers. The last step should be a person or group that will pick up, not a channel nobody watches at night.
- Timeouts match the stakes. Five minutes is reasonable for an outage that costs money every minute. Longer is a choice to wait.
- Urgency decides the route. Critical and non-critical alerts from the same service rarely deserve the same policy. Splitting them is the fastest way to cut night pages without missing real incidents.
- It pages through more than one channel. A phone call wakes people; a chat message does not. Use both for the steps that matter. Which channel reaches a sleeping person is the subject of which channel wakes an engineer, and an AI can place the call too: see how voice paging works.
- It is short. A long chain hides the question of who is really responsible. Three or four steps is usually enough.
Common shapes
Most policies are a variation on one of four shapes:
- Primary, then backup, then a manager. The classic chain, with a wait between steps. It suits small and mid-size teams with one rotation.
- Channel first, person after silence. The alert posts to the team channel, and only pages someone if nobody reacts within a set time. It suits warnings that rarely need a human at night.
- Different route by time of day. Working hours go to the channel and the whole team; nights go to the on-call engineer alone. Only alerts that are truly urgent are allowed to page at night.
- Different route by region. Each region’s schedule is on call for its own daytime, so nobody is woken when a colleague on the other side of the world is at their desk.
Pick the simplest shape that covers your risk. A team of five rarely needs more than the first, with a second policy for anything that should wait until morning.
Designing the rules as a team
An escalation policy encodes a decision the team makes together: who gets woken for which problem, and how quickly the next person steps in. Teams that design it well agree on a few rules before they write any steps:
- One owner per policy. The team whose service it covers reviews it and changes it; nobody edits another team’s policy in the middle of the night.
- Routing and escalation stay separate. Routing rules decide which policy an alert reaches, by service, severity or source. The policy decides who is paged. Keeping the two apart keeps each one simple to reason about.
- Schedules, not names. Steps point at on-call schedules, so the policy survives holidays, handoffs and people changing teams.
- A reason for every timeout. When the team writes down why a step waits five minutes and not fifteen, a later edit changes it on purpose rather than by accident.
The design also decides how the rotation feels. A policy that pages the same person first for everything turns one engineer into the team’s alarm clock. Spreading first-line duty across a rotation, and keeping the number of first-line alerts low, is a design choice as much as a staffing one.
How to test a policy
A policy that has never fired is a guess. Before relying on one, page it on purpose, let a step time out, and read the schedule behind it. The drill in escalation policy anti-patterns gives the full checklist, and after a real incident the question is whether the policy paged the right person at the right time.
The workflow around the policy
The policy is only the first part of a workflow. Once someone acknowledges, escalation stops and that person owns the alert: they can reassign it, bring in another team, or escalate by hand when they need help. Who coordinates the wider response is described in the incident commander in a small company. How fast that happens is what MTTA and MTTR measure, and the pay and fairness of the rotation behind it is a separate question the team should settle in advance: see on-call compensation.
Frequently asked questions
What is the difference between an escalation policy and an on-call schedule?
A schedule says who is on call and when. An escalation policy says what happens when an alert arrives: whom to page first, how, and what to do if that person does not answer. A policy usually points at a schedule, so it pages whoever is on call at that moment.
How long should each step wait before escalating?
Long enough for a person to be woken, reach a screen and acknowledge, and short enough that the cost of waiting stays acceptable. Five minutes is a common starting point for urgent alerts. Set the wait from the stakes of the service, then adjust it after incidents where the page went unanswered.
Should every alert use the same escalation policy?
No. Urgent alerts and low-priority ones deserve different routes. An outage on a payment flow can wake the whole chain, while a slow-burning warning can wait for morning in a channel. Splitting policies by urgency is the quickest way to cut night pages without missing real incidents.
What happens if nobody acknowledges the alert?
The policy moves to its next step, and when it runs out of steps it either repeats from the top or stops, depending on how it is set. A policy that stops leaves the alert unanswered, so most teams end the chain with a person or group who will always respond, and repeat it.
How often should we review our escalation policies?
Review them after every incident where the page went unanswered or reached the wrong person, and whenever the team, its services or its rotation change. A short review on a set schedule catches the quiet drift, such as a step that still points at someone who has left the team.