Moving your on-call to a new tool without a missed page

Published Updated 8 min read

To move on-call to a new tool without a missed page, inventory what pages whom today, map schedules, policies and alert sources, run both tools side by side, and test with drills. Then cut over one team at a time with a rollback ready, and keep the old tool until the new one has proved itself.

Where pages get missed

A migration is risky because the new setup has to reproduce, on day one, everything the old one did for years, including the parts nobody wrote down. Missed pages in a move come from a short list of causes: an alert source that was forgotten, a contact method that was never verified, a gap in a schedule that the import did not reproduce, or a routing rule that was copied wrongly. None of these is exotic. All of them can be found before the cutover by looking in the right places.

For the business, the stakes are plain. A page that nobody receives is an outage that nobody fixes, and customers find out first. The plan below is slower than a weekend switch. It is built so that nothing depends on luck, and that matters more than speed.

Step 1: inventory what pages whom today

Start by listing what exists. The aim is a complete picture of how an alert travels from a monitor to a person.

ItemWhat to recordWhere it hides
SchedulesLayers, rotations, handover times, time zonesOverrides and one-off swaps
Escalation policiesSteps, waits, who or what each step pagesPolicies for one service nobody touches
Services and routing rulesWhich alerts reach which policyCatch-all rules and severity filters
Alert sourcesEvery monitor, integration and webhook that sendsScripts, cron jobs, email forwards, old keys
People and contact methodsPhone, SMS, app, chat for each personNumbers that are out of date
Notification rulesHow and when each person wants to be reachedPersonal settings that override the defaults
AutomationAnything that calls the tool’s APIDeploy scripts and status tooling

Whether you are leaving PagerDuty, Opsgenie or another tool, export what you can, from the tool’s export or its API, so you have a record to check the new setup against. The list of alert sources is the one to spend most time on. Ask every team which systems send alerts, then check their configuration rather than relying on memory.

Step 2: decide what moves

Not everything deserves to move. Retire services nobody runs, policies nobody has triggered in a long time, and alerts that the team has already decided to delete. A migration is a chance to clean up, and it is easier to do before the move than after. The work is described in the 30-day alert fatigue plan, and the mistakes to look for in policies are collected in escalation policy anti-patterns.

Write down what you are leaving behind and why. A list of decisions beats a silent omission, because someone will ask later why an alert stopped.

Step 3: map old to new

Terms differ between tools, so build a short translation table, and decide how each concept carries across.

In the old toolIn the new toolDecide
Schedule and layersScheduleTime zones and handover times
Escalation policyEscalation policyWaits and what each step pages
ServiceServiceOwnership and runbook links
Integration or routing keyAlert source or webhookWhich alerts, and to which service
Urgency or prioritySeverityHow the levels line up

Check the small things too: a time zone that differs by one hour can shift a handover for a whole region. The new tool’s own list of integrations shows which sources it accepts, so check that every source on your list has a route.

Step 4: run both side by side

The safest way to test a new setup is to send it real alerts while the old one keeps paging. Point every alert source at both tools. Set the new tool’s paging steps to a test channel at first, so nobody is paged twice. Then compare, daily at the start:

  • Did every alert that reached the old tool also reach the new one?
  • Was it routed to the same service and policy?
  • Would the right person have been paged, by the right channel, at the right time?

Keep this up for a full rotation, including a weekend and a handover. Fix each difference as it appears. The aim is an ordinary week in which the two tools agree, so that the cutover changes only who sends the page.

Step 5: test with drills

Reading a configuration is not a test. Run drills with real people. Page each person through every channel their steps use, as described in which channel wakes an engineer, and confirm they can acknowledge. Let a step time out and watch the next one fire. Check that overrides and holiday cover are honoured, and that the handover hour is right for each time zone. Include a failure: what happens if the new tool cannot send? Write down what the team does then.

Step 6: cut over in order

Switch in small, reversible steps. A sequence that works for most teams:

StepActionWhy
1Freeze changes to schedules and policiesSo the two do not drift while you switch
2Pick a low-risk team or service to go firstA small blast radius if something is wrong
3Point its paging steps at the people, not the test channelThe first real pages come through the new tool
4Turn off paging in the old tool for that team onlySo nobody is paged twice
5Watch the first real alerts, together, for a few daysConfirm it behaves as in the drills
6Move the next team, then the nextOne change at a time, with the rollback ready

Choose the moment with care: in working hours, in the middle of the week, when people are present and not just before a holiday.

Incidents still open at cutover. Avoid switching a team while one of its incidents is running. If one is open, let it finish in the old tool, or hand its alerts and follow-ups over explicitly, and check that anything acknowledged there has an owner in the new tool. An alert that stays acknowledged in a system nobody watches is the quietest way to lose a page.

Step 7: a rollback you can use

Write the rollback before the cutover. It should say what triggers it, such as a missed page or an alert with no route, who decides, and the exact steps to point sources back at the old tool. Leave the old configuration intact for as long as the rollback must stay possible. A rollback that has to be invented under pressure is not a rollback. If an incident is under way while you roll back, the incident commander decides, as with any other change.

Step 8: tell the team

People carry the pager, so they need to know what is changing and when. Send the dates early, then a one-page cheat sheet: how to acknowledge, where alerts arrive, how to hand over, and who to contact if something looks wrong. Ask everyone to install the tool’s app, if it has one, verify their contact details and allow the paging numbers through do-not-disturb, then check that they have. Update runbooks that mention the old tool’s name or links.

After the cutover, keep the old tool for a quiet period, export the history you need, and then switch it off in stages. Running two tools for a time is a real cost. It is small next to a missed page, and it is the reason the migration did not cost a customer anything.

Frequently asked questions

How long should we run both tools side by side?

Long enough to cover a full rotation, including a weekend and at least one handover, so that every schedule and policy has been exercised on real alerts. Longer is safer, and the cost of running two tools for a while is small beside the cost of a missed page.

What is the most common cause of a missed page in a migration?

An alert source that nobody remembered. A script, a scheduled job or an old integration sends alerts by a route that was never listed, and it keeps pointing at the old tool after the cutover. The inventory of alert sources is the step that prevents it.

Should we clean up alerts before moving?

Yes, where you can. Moving noisy alerts and stale policies into a new tool carries the problems over. A short clean-up first means you migrate what is in use and worth keeping, and the team is not tested on the old noise during the cutover.

What should a rollback look like?

A written list of steps that points alert sources back at the old tool, with named people, a clear trigger such as a missed page, and the old configuration left intact for as long as the rollback stays possible. Agree it before the cutover, not during it.

When is the old tool safe to switch off?

After a quiet period in which every source has been sending to the new tool, every team has received and answered real pages there, and nothing has reached the old one. Export the history you need for records first, then switch it off in stages, not all at once.