Escalation policy anti-patterns, and how to test yours
The common escalation policy anti-patterns are a single point of failure, timeouts that are too long or too short, paging everyone at once, stale schedules, and no backup for the backup. Find them by testing: page the policy on purpose, let each step time out, and read the schedule behind it.
Why anti-patterns hide
An escalation policy rarely fails on a normal night. It fails on the night when the primary’s phone is dead, the schedule has a gap or the alert is bigger than anyone planned for. Until then it looks fine, which is why its faults survive. The policy is a design, and like any design it has failure modes that only appear under stress. Naming the common ones makes them easier to see. The basics of how a policy works are in what an escalation policy is.
The anti-patterns at a glance
| Anti-pattern | What goes wrong | The fix |
|---|---|---|
| A single point of failure | One person or channel is the whole path | A second person and a second channel at each step |
| Timeouts that are too long | An outage runs while nobody answers | Set the wait from the stakes of the service |
| Timeouts that are too short | The next person is paged before the first can respond | Allow time to wake and acknowledge |
| Paging everyone at once | Everyone assumes someone else will act | One step at a time, with a clear first responder |
| Stale schedules | Pages go to people who left, or to nobody | Regular schedule review, tied to joiners and leavers |
| No backup for the backup | The chain ends at a channel nobody watches | A last step that always reaches a person who answers |
| Acknowledge and vanish | An incident is acknowledged and then forgotten | Re-notify until resolved or updated |
| One policy for every alert | Minor alerts wake people, or major ones wait | Separate policies by urgency |
| A policy nobody owns | It drifts and nobody fixes it | One owning team, named in the policy |
Who is paged
A single point of failure. Look for the step that has one person, one channel or one number. A policy that pages the primary by push notification only is as reliable as that one app on that one phone. Give every important step two ways to reach a person, and the chain a second person behind the first. How the channels compare is set out in which channel wakes an engineer.
Everyone at once. It feels safe to page the whole team. In practice it spreads responsibility so thin that no one takes it, and it wakes many people for a problem that needs one. Page one person, wait, then the next. If the alert truly needs everyone, that is a different, rare policy.
No backup for the backup. Chains often end at a shared channel or a manager who is on holiday. The last step should be someone or something that will answer, and the policy should repeat if it reaches the end with no response.
When they are paged
Timeouts are the least examined setting in most policies. Too long, and an outage continues in silence. A step that waits half an hour on an urgent alert is a decision to lose half an hour. Too short, and the second person is woken while the first is still finding their laptop, which creates two responders and confusion.
Set each wait from the cost of delay and the time a person needs. The number matters less than the reason: write down why a step waits as long as it does, so a later change is deliberate. Then check the numbers against real incidents in MTTA and MTTR, where a long wait shows up as a long time to acknowledge.
Acknowledge and vanish. Acknowledging stops escalation. It does not mean anyone is working. Add a rule that re-notifies if the alert is not resolved or updated within a set time, so a half-asleep acknowledgement cannot leave an incident alone.
Whether it is current
Stale schedules. People join, leave, change teams and go on leave. A schedule that has not been reviewed will page a person who left, or contain a gap where nobody is on call. Tie the review to events: a joiner, a leaver, a change of team. Look eight weeks ahead for gaps on a regular basis.
One policy for everything. A policy that treats a slow-burning warning like an outage will either wake people for nothing or delay the outage. Split policies by urgency. Routing decides which policy an alert reaches, and the policy decides who is paged. Cutting the volume of what reaches the urgent policy is the work of an alert fatigue plan.
Nobody owns it. Give each policy an owning team, written in it. Owners review after incidents and on a schedule. Without an owner, the rules drift, and the drift is only noticed on the bad night.
A worked example
This audit is invented. A team reviews its critical policy and finds four faults in half an hour. The first step pages the primary by push notification only. The second step’s timeout is half an hour. The last step is the team channel. The schedule has a gap of one weekend, three weeks ahead, when a colleague is on leave.
Each fault has a small fix. Add SMS and a call to the first step, and shorten the timeout to a few minutes, as urgent alerts need. Add the engineering manager on call as the last step, and repeat the chain. Assign the weekend to a named engineer. Written this way, the audit takes an afternoon and removes four ways for a page to be lost. The team also adds one rule: any change to the schedule triggers a test page.
A drill to test yours
A drill finds what reading cannot. Run it in a quiet period, tell the team, and use a test alert that looks like a real one. A step passes when the page breaks through, or when the next step fires on time. Silence is the only failure.
- Page the first step. Confirm the right person is reached on every channel the step uses, and that they can acknowledge.
- Do not acknowledge. Let the step time out. Confirm the second step pages when it should, not sooner or later.
- Make the primary unreachable. Ask them to put their phone in a focus mode. See whether the path still gets through.
- Reach the end. Confirm the last step reaches a person who answers, and that the policy repeats.
- Acknowledge and leave. Check that the re-notification fires if nobody updates the alert.
- Read the schedule. Look for gaps in the next eight weeks, and for anyone on it who should not be.
- Test at a realistic hour. A test at noon proves little about three in the morning.
- Write down what you found. Fix each fault, assign an owner, and set the date of the next drill.
Repeat the drill after any change, and after any incident where a page went unanswered.
Frequently asked questions
What is the most common escalation policy mistake?
A chain that depends on one person or one channel. It works on every night when that person is reachable and fails on the night they are not. The remedy is a second person at every step that matters, a second channel for each person, and a last step that always reaches someone who will answer.
Is it bad to page everyone at once?
For most alerts, yes. When everyone is paged, each person assumes someone else will respond, and the whole team loses sleep for one problem. Reserve it for the rare alert that truly needs the whole team, and escalate one step at a time for the rest.
How do we know a timeout is right?
Ask what a person needs to wake, reach a screen and acknowledge, and what the cost is of waiting longer. Too short, and the next person is paged while the first is still getting out of bed. Too long, and an outage runs while nobody responds. Adjust after real incidents.
Does acknowledging an alert mean it is being fixed?
No. Acknowledging only stops the escalation. A re-notification rule, described above, stops an acknowledgement without action from leaving an incident unattended.
How often should we test an escalation policy?
After any change to the policy, the schedule or the people on it, after every incident where a page went unanswered, and on a regular schedule such as monthly or quarterly. A short drill takes little time and finds problems while they are cheap to fix.