Moving your on-call to a new tool without a missed page
To move on-call to a new tool without a missed page, inventory what pages whom today, map schedules, policies and alert sources, run both tools side by side, and test with drills. Then cut over one team at a time with a rollback ready, and keep the old tool until the new one has proved itself.
Where pages get missed
A migration is risky because the new setup has to reproduce, on day one, everything the old one did for years, including the parts nobody wrote down. Missed pages in a move come from a short list of causes: an alert source that was forgotten, a contact method that was never verified, a gap in a schedule that the import did not reproduce, or a routing rule that was copied wrongly. None of these is exotic. All of them can be found before the cutover by looking in the right places.
For the business, the stakes are plain. A page that nobody receives is an outage that nobody fixes, and customers find out first. The plan below is slower than a weekend switch. It is built so that nothing depends on luck, and that matters more than speed.
Step 1: inventory what pages whom today
Start by listing what exists. The aim is a complete picture of how an alert travels from a monitor to a person.
| Item | What to record | Where it hides |
|---|---|---|
| Schedules | Layers, rotations, handover times, time zones | Overrides and one-off swaps |
| Escalation policies | Steps, waits, who or what each step pages | Policies for one service nobody touches |
| Services and routing rules | Which alerts reach which policy | Catch-all rules and severity filters |
| Alert sources | Every monitor, integration and webhook that sends | Scripts, cron jobs, email forwards, old keys |
| People and contact methods | Phone, SMS, app, chat for each person | Numbers that are out of date |
| Notification rules | How and when each person wants to be reached | Personal settings that override the defaults |
| Automation | Anything that calls the tool’s API | Deploy scripts and status tooling |
Whether you are leaving PagerDuty, Opsgenie or another tool, export what you can, from the tool’s export or its API, so you have a record to check the new setup against. The list of alert sources is the one to spend most time on. Ask every team which systems send alerts, then check their configuration rather than relying on memory.
Step 2: decide what moves
Not everything deserves to move. Retire services nobody runs, policies nobody has triggered in a long time, and alerts that the team has already decided to delete. A migration is a chance to clean up, and it is easier to do before the move than after. The work is described in the 30-day alert fatigue plan, and the mistakes to look for in policies are collected in escalation policy anti-patterns.
Write down what you are leaving behind and why. A list of decisions beats a silent omission, because someone will ask later why an alert stopped.
Step 3: map old to new
Terms differ between tools, so build a short translation table, and decide how each concept carries across.
| In the old tool | In the new tool | Decide |
|---|---|---|
| Schedule and layers | Schedule | Time zones and handover times |
| Escalation policy | Escalation policy | Waits and what each step pages |
| Service | Service | Ownership and runbook links |
| Integration or routing key | Alert source or webhook | Which alerts, and to which service |
| Urgency or priority | Severity | How the levels line up |
Check the small things too: a time zone that differs by one hour can shift a handover for a whole region. The new tool’s own list of integrations shows which sources it accepts, so check that every source on your list has a route.
Step 4: run both side by side
The safest way to test a new setup is to send it real alerts while the old one keeps paging. Point every alert source at both tools. Set the new tool’s paging steps to a test channel at first, so nobody is paged twice. Then compare, daily at the start:
- Did every alert that reached the old tool also reach the new one?
- Was it routed to the same service and policy?
- Would the right person have been paged, by the right channel, at the right time?
Keep this up for a full rotation, including a weekend and a handover. Fix each difference as it appears. The aim is an ordinary week in which the two tools agree, so that the cutover changes only who sends the page.
Step 5: test with drills
Reading a configuration is not a test. Run drills with real people. Page each person through every channel their steps use, as described in which channel wakes an engineer, and confirm they can acknowledge. Let a step time out and watch the next one fire. Check that overrides and holiday cover are honoured, and that the handover hour is right for each time zone. Include a failure: what happens if the new tool cannot send? Write down what the team does then.
Step 6: cut over in order
Switch in small, reversible steps. A sequence that works for most teams:
| Step | Action | Why |
|---|---|---|
| 1 | Freeze changes to schedules and policies | So the two do not drift while you switch |
| 2 | Pick a low-risk team or service to go first | A small blast radius if something is wrong |
| 3 | Point its paging steps at the people, not the test channel | The first real pages come through the new tool |
| 4 | Turn off paging in the old tool for that team only | So nobody is paged twice |
| 5 | Watch the first real alerts, together, for a few days | Confirm it behaves as in the drills |
| 6 | Move the next team, then the next | One change at a time, with the rollback ready |
Choose the moment with care: in working hours, in the middle of the week, when people are present and not just before a holiday.
Incidents still open at cutover. Avoid switching a team while one of its incidents is running. If one is open, let it finish in the old tool, or hand its alerts and follow-ups over explicitly, and check that anything acknowledged there has an owner in the new tool. An alert that stays acknowledged in a system nobody watches is the quietest way to lose a page.
Step 7: a rollback you can use
Write the rollback before the cutover. It should say what triggers it, such as a missed page or an alert with no route, who decides, and the exact steps to point sources back at the old tool. Leave the old configuration intact for as long as the rollback must stay possible. A rollback that has to be invented under pressure is not a rollback. If an incident is under way while you roll back, the incident commander decides, as with any other change.
Step 8: tell the team
People carry the pager, so they need to know what is changing and when. Send the dates early, then a one-page cheat sheet: how to acknowledge, where alerts arrive, how to hand over, and who to contact if something looks wrong. Ask everyone to install the tool’s app, if it has one, verify their contact details and allow the paging numbers through do-not-disturb, then check that they have. Update runbooks that mention the old tool’s name or links.
After the cutover, keep the old tool for a quiet period, export the history you need, and then switch it off in stages. Running two tools for a time is a real cost. It is small next to a missed page, and it is the reason the migration did not cost a customer anything.
Frequently asked questions
How long should we run both tools side by side?
Long enough to cover a full rotation, including a weekend and at least one handover, so that every schedule and policy has been exercised on real alerts. Longer is safer, and the cost of running two tools for a while is small beside the cost of a missed page.
What is the most common cause of a missed page in a migration?
An alert source that nobody remembered. A script, a scheduled job or an old integration sends alerts by a route that was never listed, and it keeps pointing at the old tool after the cutover. The inventory of alert sources is the step that prevents it.
Should we clean up alerts before moving?
Yes, where you can. Moving noisy alerts and stale policies into a new tool carries the problems over. A short clean-up first means you migrate what is in use and worth keeping, and the team is not tested on the old noise during the cutover.
What should a rollback look like?
A written list of steps that points alert sources back at the old tool, with named people, a clear trigger such as a missed page, and the old configuration left intact for as long as the rollback stays possible. Agree it before the cutover, not during it.
When is the old tool safe to switch off?
After a quiet period in which every source has been sending to the new tool, every team has received and answered real pages there, and nothing has reached the old one. Export the history you need for records first, then switch it off in stages, not all at once.