Suppress by policy: alert rules a human can defend
Suppress alerts by policy when a written rule, not a habit, decides that nobody is paged. A defensible rule says what it matches, why holding it back is safe, who owns it and when it ends, and it leaves the alert visible for review. If you cannot explain a rule to a colleague at 3 a.m., do not suppress.
Habit versus policy
Most suppression starts as a habit. Someone mutes an alert during an incident, a filter is added quickly, a threshold is nudged. Each action makes sense at the time. Months later there is a collection of silences that nobody owns, nobody can explain and nobody dares remove, and one of them is hiding something that matters.
Suppression by policy is the alternative. The decision is a written rule, made on purpose, with a reason that a colleague can read and dispute. The aim is not fewer rules. It is rules that can be defended: to a teammate, to a manager, and to yourself in the middle of the night. The different kinds of noise reduction are set out in deduplication, correlation and suppression. This page is about the third kind.
What a defensible rule contains
A rule that can be defended has six parts.
| Part | The question it answers | Example, invented |
|---|---|---|
| Match | Exactly which alerts does it cover? | Alert “disk usage” on the batch hosts, warning level |
| Reason | Why is holding it back safe? | It fires nightly during a known job and clears alone |
| Owner | Who answers for this rule? | The data platform team |
| Scope | Where does it apply, and where does it not? | Batch hosts only, not the database hosts |
| Expiry | When is it reviewed or removed? | Reviewed monthly, expires at the end of the quarter |
| Visibility | Where can a person see what it held back? | A daily list, linked from the team channel |
If any row is empty, the rule is a habit with a name. A missing reason is the most common gap, and the most serious.
Rules that say why
The reason decides how safe a rule is. A few reasons hold up, and a few do not.
| Reason | Safe when | Unsafe when |
|---|---|---|
| Planned maintenance | The window is stated and ends on its own | The window is open-ended |
| Consequence of a known failure | The upstream cause is firing and being handled | The dependency map is out of date |
| Low priority out of hours | The alert can wait until working hours | Nobody looks at it in working hours either |
| Long history of clearing with no action | A score and its reasons show it, and the alert has an owner | The history is short, or the alert is new |
| It is annoying | Never: fix or delete the alert | Always |
The last row is the test. If the only reason is annoyance, the alert is broken, and the fix is to repair or remove it, not to hide it.
Noise scores and their reasons
A noise score, sometimes called a flakiness score, is a measure of how likely an alert is to be noise. A tool builds it from data the system already has: how often the alert fires, how quickly it clears, and whether anyone ever acts on it. A high score means the alert usually fires and disappears without help.
A score is useful because it gives a rule something objective to stand on: “alerts above this score stay off this policy”. Some tools let an escalation policy carry exactly that kind of condition, so the alert stays on record but does not page through that policy. Check what happens next: if the tool uses the first policy that matches, an alert that fails one policy’s condition can fall through to another policy that still pages. It is only defensible if the reasons are visible. “Score high” is a number. “Fired often this week, cleared alone each time, no acknowledgement” is an argument a person can check.
The limits deserve stating. A score is built from history, so a new alert has none. A flaky alert can still be the first sign of a real failure. And a score can drift when the system changes. Treat it as evidence for a rule that a person still owns.
Reviewing what was held back
A suppression rule is only as safe as its review. Three habits keep it honest.
- A weekly list. Look at what each rule held back. Anything surprising is a reason to narrow or end the rule.
- A post-incident question. After each incident, ask whether a suppressed alert was the earliest sign. If it was, the rule needs changing, and the answer belongs in the review.
- Expiry. A rule that ends on its own has to be renewed on purpose. That is the cheapest way to clear out rules that no longer serve.
Keep a short change log, too: who added or changed a rule, when, and why. It makes the rules something a team can read, which is what separates policy from habit.
Test a rule before it goes live
The best test uses your own history. Take the last month of alerts and apply the new rule as a dry run: what would it have held back? Read that list. If everything on it is something you are glad not to have been paged for, the rule is ready. If one item makes you wince, narrow the match or drop the rule.
Keep a protected class that no rule can suppress. Alerts that signal customer-facing failure or a security event can be grouped and deduplicated, but they should always reach a person. Changing the protected list should need more than one person’s agreement.
A rule review, worked through
This example is invented. A team proposes to suppress a “queue depth high” alert on its notification service, because it fires most evenings.
The reviewer asks the questions the table above lists. The match is one alert on one service, which is narrow enough. The reason is that the queue backs up during a nightly bulk send and drains within minutes. The owner is the notifications team. The dry run over the last month shows the alert fired on most evenings and cleared on its own each time, with one exception: a night when the queue kept growing because a worker had crashed.
That exception decides the outcome. Suppressing the alert would have hidden the one occurrence that mattered. So the team changes the rule: hold the alert back while the queue drains within a set time, and let it page if the depth is still high after that. It expires in a month, and the held-back list goes to the team channel each morning. The rule is narrower, has a reason and an owner, and would have caught the crashed worker.
Mistakes to avoid
- A broad match. “All warnings from this team” is a habit, not a rule.
- No end. A rule without an expiry outlives the reason for it.
- Hiding instead of fixing. If the alert is wrong, correct the alert.
- Silent rules. If nobody sees what was held back, nobody can tell when the rule has gone wrong.
- Too many owners, or none. One named owner per rule, always.
Where the rules live
Suppression belongs next to the decision about who is told. An escalation policy already answers “who is paged for this alert and when”, so conditions on it are a natural home for “and not for these”. Whatever the tool, keep rules in one place, version them if you can, and give them the same care as production changes. When the noise is a general flood rather than a few known alerts, start with the 30-day plan for alert fatigue, and use suppression for what is left.
Frequently asked questions
What is alert suppression by policy?
It is a written rule that stops a class of alerts from notifying anyone, with a stated reason, an owner and an end or review date. The alert is still recorded. The difference from muting is that the decision is made once, on purpose, and can be read and challenged later, instead of living in someone's head.
When is it safe to suppress an alert?
When you can say why nobody needs to be told. Examples are a planned maintenance window, an alert that is a known consequence of a failure already being handled, or an alert that has a long record of firing and clearing with no action. If the reason is only that the alert is annoying, fix or delete the alert.
What is a noise score?
It is a measure of how likely an alert is to be noise, worked out from its history: how often it fires, how quickly it clears, and whether anyone ever acts on it. It is evidence for a rule, not a rule. Show the reasons behind the score so a person can check it.
How do we make sure suppressed alerts do not hide a real incident?
Keep them visible in a list, review that list each week, and after every incident ask whether a suppressed alert was the earliest sign. Give every rule an expiry, and test new rules against past alerts to see what they would have held back.
Should any alerts never be suppressed?
Yes. Agree a protected class in advance, such as alerts that signal customer-facing failure or a security event. They can be grouped and deduplicated, but no rule may hold them back from a person.