Alert fatigue: a 30-day plan
Alert fatigue is what happens when people get more alerts than they can act on, so they stop trusting any of them. A 30-day plan fixes it in order: audit what fires, sort each alert as actionable or informational, delete or reroute the rest, tune what remains, and review every week. Give it an owner, or nothing changes.
Alert fatigue is a reliability problem
Alert fatigue is usually described as a morale problem: people are tired of being woken. It is also a reliability problem, and the second is the reason to act. When most alerts need no action, people learn that an alert usually means nothing. They acknowledge to silence the noise, skim messages, and answer more slowly. Then a real incident arrives looking like every other alert, and it is treated the same way.
The effect shows up in the numbers a team already tracks. Time to acknowledge, described in MTTA vs MTTR, can move either way: slower acknowledgements, or suspiciously fast ones with no follow-up action. Either is often the first sign. The fix is not to tell people to try harder. It is to change the alerts, and thirty days is enough to make a visible difference if the work is done in order.
The plan at a glance
| Week | Goal | What you have at the end |
|---|---|---|
| 1 | Audit what fires and who acts | A ranked list of alert rules with facts beside each |
| 2 | Sort into actionable, informational, noise | Noise deleted, informational alerts rerouted |
| 3 | Tune what is left and fix causes | Fewer, clearer pages with a runbook each |
| 4 | Route by owner and set the review | Every alert has an owner, and a weekly review runs |
The order matters. Tuning before you have sorted means you tune alerts that should not exist, and deleting before you have audited means arguing from feeling instead of facts.
Week 1: audit
Collect the last month of alerts, or as much history as you have. For each alert rule, record how often it fired, who was notified, how long it took to acknowledge, and, most importantly, what the person did about it. That last column is the one people skip and the one that matters.
A simple sheet is enough. Sort it by volume. A small number of rules usually produces a large share of the pages, so the top of the list is where the effort pays off. Add one question per rule: “If this fired at 3 a.m., what should the person do?” Many rules have no answer, and those are the first candidates.
An illustrative slice of an audit sheet, with invented rules, shows how the verdicts follow from the facts:
| Alert rule | Fired often | Who acted, and how | Verdict |
|---|---|---|---|
| Checkout error rate is high | Rarely | On-call rolled back or fixed | Actionable: keep as a page |
| CPU is high on a web instance | Daily | Nobody; it clears alone | Noise: delete, or keep as a graph |
| Disk is nearly full, batch host | Weekly | Someone clears space next morning | Informational: ticket, working hours |
| Queue depth is above the limit | Often | Restart, then it comes back | Actionable, but fix the cause |
Include the people. Ask the current on-call engineers which alerts they mute, ignore or dread. They already know, and their list usually matches what the data shows.
Week 2: classify, then delete or reroute
Sort every rule into one of three classes.
| Class | The test | What happens to it |
|---|---|---|
| Actionable | A person must do something now, and can | Stays a page, with an owner and a runbook |
| Informational | Useful to know, not urgent | Moves to a channel, a ticket or a daily summary |
| Noise | Nobody acts on it, or it fixes itself | Deleted, with a note of why |
The test for a page is simple: it needs a human, it needs one now, and the human can do something about it. Anything that fails one of the three does not belong in the middle of the night.
If an alert is real but should not page anyone, a written suppression rule is safer than a mute: see suppress by policy.
Be careful when deleting. Turn a rule off first, keep its definition for a few weeks, and note the reason in writing. If nothing goes wrong, remove it. A recorded reason is also what lets a nervous colleague accept the change.
Week 3: tune and fix causes
What remains needs to be made better. Four moves cover most cases.
- Page on symptoms. Alert on what customers feel, such as failed requests or slow pages, and treat low-level causes such as high CPU as information unless they lead to a symptom.
- Add a duration. An alert that must stay true for a few minutes before firing ignores blips that resolve alone.
- Group related alerts. Twenty alerts from one failure should arrive as one incident, not twenty pages. How deduplication, correlation and suppression differ tells you which to use.
- Fix the source. A flaky alert is often pointing at a real defect in the service or the check. Repairing that is better than silencing it.
Every alert that stays a page gets a runbook link, so the person woken up knows what to do first.
Week 4: route by owner and set the review
An alert without an owner has no one to fix it. Give every rule an owning team, and route it to that team’s escalation policy, with urgency deciding how far it can escalate. Urgent alerts can wake people. The rest should wait for working hours.
Protect the people doing the work, too: the signs that a rotation is wearing out are covered in on-call burnout. Then set the review.
Thirty minutes each week is enough: look at the noisiest alerts, the ones nobody acted on, and any new alerts added since the last meeting. Assign each problem to a person. The review is what turns a one-off clean-up into a habit.
Measure without vanity numbers
The tempting number is the count of alerts deleted. It says how busy the team was, not whether anyone is better off. Choose measures that follow the people:
- Pages per person per shift, especially at night, tracked over time.
- The share of pages that led to action. Describe it in words to the team: “most pages now need us”.
- The noisiest five rules, and whether they are the same five as last month.
- Time to acknowledge, by severity.
Look at the trend across a few weeks. One quiet week can be luck.
Objections you will hear
- “What if we delete something that mattered?” Turn rules off before deleting, keep the definition and the reason, and put borderline alerts in a channel first. The rule can be restored in a minute. Watch that service more closely while it is off.
- “The alert has always been there.” Age is not a reason. Ask what the person does when it fires.
- “We cannot afford the time.” The time is already being spent, in interrupted sleep and slow responses. The plan moves it to a place where it pays back.
- “The customer contract needs it.” Then find the action the contract requires. If that is a report, route it to one, not to a pager.
Who owns it
Alert fatigue belongs to the team that owns the service, and it needs a manager’s backing. Engineers know which alerts are useless. What they often lack is permission and time to delete them. A team rule that any alert can be challenged with a question, “what would you do?”, makes the conversation short. A manager who protects time for the clean-up and thirty minutes a week for the review, and who accepts a temporary rise in risk while noise is removed, makes the plan work.
Set a rule for new alerts, too: an owner, a clear action, a runbook link and an urgency, or it does not ship. That single rule, applied consistently, prevents most of the return of the noise.
Frequently asked questions
What is alert fatigue?
It is the gradual loss of attention and trust that follows a steady flow of alerts that do not need action. People start to skim, silence or acknowledge without reading. The danger is not annoyance. It is that the one alert that matters looks the same as the hundred before it.
How many alerts should a team get?
There is no universal number, because it depends on the system and the team. A useful test is whether each page needs a person to step in immediately. If a rotation regularly gets pages nobody acts on, or pages that wake people for things that could wait, the volume is too high for that team.
Should we just raise the thresholds?
Sometimes, but not first. A higher threshold hides the problem the alert was pointing at, and it can hide a real one. Start by asking what the person should do when the alert fires. If the answer is nothing, remove or reroute it. Tune a threshold only for an alert that needs action.
How do we stop the noise coming back?
Put a rule on new alerts: each needs an owner, a clear action, a runbook link and an urgency. Review the noisiest alerts every week. Noise returns when nobody owns the list, so make owning it a named job on the team.
Who is responsible for fixing alert fatigue?
The team that owns the service, with backing from whoever manages the team. Engineers know which alerts are useless, but they need permission and time to delete them. A manager who protects that time makes the plan work; one who only asks for fewer alerts does not.