Alert fatigue: a 30-day plan

Published Updated 8 min read

Alert fatigue is what happens when people get more alerts than they can act on, so they stop trusting any of them. A 30-day plan fixes it in order: audit what fires, sort each alert as actionable or informational, delete or reroute the rest, tune what remains, and review every week. Give it an owner, or nothing changes.

Alert fatigue is a reliability problem

Alert fatigue is usually described as a morale problem: people are tired of being woken. It is also a reliability problem, and the second is the reason to act. When most alerts need no action, people learn that an alert usually means nothing. They acknowledge to silence the noise, skim messages, and answer more slowly. Then a real incident arrives looking like every other alert, and it is treated the same way.

The effect shows up in the numbers a team already tracks. Time to acknowledge, described in MTTA vs MTTR, can move either way: slower acknowledgements, or suspiciously fast ones with no follow-up action. Either is often the first sign. The fix is not to tell people to try harder. It is to change the alerts, and thirty days is enough to make a visible difference if the work is done in order.

The plan at a glance

WeekGoalWhat you have at the end
1Audit what fires and who actsA ranked list of alert rules with facts beside each
2Sort into actionable, informational, noiseNoise deleted, informational alerts rerouted
3Tune what is left and fix causesFewer, clearer pages with a runbook each
4Route by owner and set the reviewEvery alert has an owner, and a weekly review runs

The order matters. Tuning before you have sorted means you tune alerts that should not exist, and deleting before you have audited means arguing from feeling instead of facts.

Week 1: audit

Collect the last month of alerts, or as much history as you have. For each alert rule, record how often it fired, who was notified, how long it took to acknowledge, and, most importantly, what the person did about it. That last column is the one people skip and the one that matters.

A simple sheet is enough. Sort it by volume. A small number of rules usually produces a large share of the pages, so the top of the list is where the effort pays off. Add one question per rule: “If this fired at 3 a.m., what should the person do?” Many rules have no answer, and those are the first candidates.

An illustrative slice of an audit sheet, with invented rules, shows how the verdicts follow from the facts:

Alert ruleFired oftenWho acted, and howVerdict
Checkout error rate is highRarelyOn-call rolled back or fixedActionable: keep as a page
CPU is high on a web instanceDailyNobody; it clears aloneNoise: delete, or keep as a graph
Disk is nearly full, batch hostWeeklySomeone clears space next morningInformational: ticket, working hours
Queue depth is above the limitOftenRestart, then it comes backActionable, but fix the cause

Include the people. Ask the current on-call engineers which alerts they mute, ignore or dread. They already know, and their list usually matches what the data shows.

Week 2: classify, then delete or reroute

Sort every rule into one of three classes.

ClassThe testWhat happens to it
ActionableA person must do something now, and canStays a page, with an owner and a runbook
InformationalUseful to know, not urgentMoves to a channel, a ticket or a daily summary
NoiseNobody acts on it, or it fixes itselfDeleted, with a note of why

The test for a page is simple: it needs a human, it needs one now, and the human can do something about it. Anything that fails one of the three does not belong in the middle of the night.

If an alert is real but should not page anyone, a written suppression rule is safer than a mute: see suppress by policy.

Be careful when deleting. Turn a rule off first, keep its definition for a few weeks, and note the reason in writing. If nothing goes wrong, remove it. A recorded reason is also what lets a nervous colleague accept the change.

Week 3: tune and fix causes

What remains needs to be made better. Four moves cover most cases.

  • Page on symptoms. Alert on what customers feel, such as failed requests or slow pages, and treat low-level causes such as high CPU as information unless they lead to a symptom.
  • Add a duration. An alert that must stay true for a few minutes before firing ignores blips that resolve alone.
  • Group related alerts. Twenty alerts from one failure should arrive as one incident, not twenty pages. How deduplication, correlation and suppression differ tells you which to use.
  • Fix the source. A flaky alert is often pointing at a real defect in the service or the check. Repairing that is better than silencing it.

Every alert that stays a page gets a runbook link, so the person woken up knows what to do first.

Week 4: route by owner and set the review

An alert without an owner has no one to fix it. Give every rule an owning team, and route it to that team’s escalation policy, with urgency deciding how far it can escalate. Urgent alerts can wake people. The rest should wait for working hours.

Protect the people doing the work, too: the signs that a rotation is wearing out are covered in on-call burnout. Then set the review.

Thirty minutes each week is enough: look at the noisiest alerts, the ones nobody acted on, and any new alerts added since the last meeting. Assign each problem to a person. The review is what turns a one-off clean-up into a habit.

Measure without vanity numbers

The tempting number is the count of alerts deleted. It says how busy the team was, not whether anyone is better off. Choose measures that follow the people:

  • Pages per person per shift, especially at night, tracked over time.
  • The share of pages that led to action. Describe it in words to the team: “most pages now need us”.
  • The noisiest five rules, and whether they are the same five as last month.
  • Time to acknowledge, by severity.

Look at the trend across a few weeks. One quiet week can be luck.

Objections you will hear

  • “What if we delete something that mattered?” Turn rules off before deleting, keep the definition and the reason, and put borderline alerts in a channel first. The rule can be restored in a minute. Watch that service more closely while it is off.
  • “The alert has always been there.” Age is not a reason. Ask what the person does when it fires.
  • “We cannot afford the time.” The time is already being spent, in interrupted sleep and slow responses. The plan moves it to a place where it pays back.
  • “The customer contract needs it.” Then find the action the contract requires. If that is a report, route it to one, not to a pager.

Who owns it

Alert fatigue belongs to the team that owns the service, and it needs a manager’s backing. Engineers know which alerts are useless. What they often lack is permission and time to delete them. A team rule that any alert can be challenged with a question, “what would you do?”, makes the conversation short. A manager who protects time for the clean-up and thirty minutes a week for the review, and who accepts a temporary rise in risk while noise is removed, makes the plan work.

Set a rule for new alerts, too: an owner, a clear action, a runbook link and an urgency, or it does not ship. That single rule, applied consistently, prevents most of the return of the noise.

Frequently asked questions

What is alert fatigue?

It is the gradual loss of attention and trust that follows a steady flow of alerts that do not need action. People start to skim, silence or acknowledge without reading. The danger is not annoyance. It is that the one alert that matters looks the same as the hundred before it.

How many alerts should a team get?

There is no universal number, because it depends on the system and the team. A useful test is whether each page needs a person to step in immediately. If a rotation regularly gets pages nobody acts on, or pages that wake people for things that could wait, the volume is too high for that team.

Should we just raise the thresholds?

Sometimes, but not first. A higher threshold hides the problem the alert was pointing at, and it can hide a real one. Start by asking what the person should do when the alert fires. If the answer is nothing, remove or reroute it. Tune a threshold only for an alert that needs action.

How do we stop the noise coming back?

Put a rule on new alerts: each needs an owner, a clear action, a runbook link and an urgency. Review the noisiest alerts every week. Noise returns when nobody owns the list, so make owning it a named job on the team.

Who is responsible for fixing alert fatigue?

The team that owns the service, with backing from whoever manages the team. Engineers know which alerts are useless, but they need permission and time to delete them. A manager who protects that time makes the plan work; one who only asks for fewer alerts does not.