Deduplication, correlation and suppression: what is the difference?

Published Updated 7 min read

Deduplication merges identical repeated alerts into one. Correlation groups different alerts that share a cause into one incident. Suppression deliberately withholds a notification by a rule. Deduplication and correlation change how alerts are presented. Suppression withholds a person's attention, so apply them in that order, and never confuse any of them with raising a threshold.

Three ways to make fewer notifications

The three terms are often used as if they meant “reduce noise”, and vendors use them loosely. They are different operations, on different data, with different risks. Getting them straight helps you pick the right one and know what it costs.

TechniqueThe question it answersWhat it does to the dataTypical failure
DeduplicationIs this the same alert as one already open?Merges copies into one, with a countA key that merges different problems
CorrelationDo these different alerts share one cause?Groups them into one incidentGrouping that hides a second, independent fault
SuppressionShould anyone be told about this alert, now?Keeps the alert, withholds the notificationA rule nobody remembers that hides a real one

The rest of this page takes each in turn.

Deduplication: the same alert, once

A check that fails every minute should not send sixty pages an hour. Deduplication recognises that the repeats are the same alert and shows one, with a count of how often it fired and when it last did.

The mechanism is a dedup key: a fingerprint built from fields that identify the alert, such as the source, the check name and the affected host or service. Two alerts with the same key are the same alert. When the key is right, nothing is lost: the count and the timestamps are still there.

It goes wrong in two ways. A key that is too broad merges different problems: the same alert name on two hosts becomes one alert, and one of the two failures is invisible. A key that is too narrow, for example one that includes a changing value such as a timestamp or a request id, never matches, and the duplicates stay. Another trap is the alert that resolves and fires again a minute later. Decide how long a resolved alert may reopen before it counts as a new one.

Correlation: different alerts, one cause

A failing database can make a dozen services alert at once: slow queries, timeouts, error rates, queue growth. Those alerts are not the same, but they are one incident. Correlation groups them so a responder sees one problem and a list of symptoms.

The signals correlation can use include:

  • Time. Alerts that start within a short window of each other.
  • Shared labels. The same service, cluster, region or team.
  • Topology. A known dependency between services, so a downstream alert joins the upstream incident.
  • Content. Similar messages or error types.
  • Past patterns. Alerts that have appeared together before.

Correlation is a judgement, and judgements can be wrong. Grouping by time alone puts unrelated failures together. Grouping that is too loose folds a second, independent fault into an incident that is already open, where nobody looks for it. Too tight, and the noise stays. The architecture that works best is the one that can explain why two alerts were grouped, lets a person split or merge an incident, and treats grouping as a suggestion, not a fact.

Suppression: a rule that withholds a notification

Suppression is different in kind. The alert exists and is recorded, but a rule decides that nobody needs to be told. Common forms:

  • Maintenance windows. Alerts from systems under planned work, for a stated period.
  • Inhibition. While an upstream failure is firing, alerts that are known consequences of it stay quiet.
  • Severity by time. Low-priority alerts wait for working hours.
  • Known-flaky alerts. An alert with a history of firing and clearing on its own is held back from paging.

Suppression is the technique that loses information for the reader, because someone decided the alert did not need attention. So it needs the most discipline: every rule has a reason, an owner and, ideally, an end date, and the suppressed alerts stay visible somewhere a person can review them. That means rules a human can defend, with the reason written down.

Not the same as raising a threshold

A threshold is part of the alert’s definition: “error rate above this value for this long”. Raising it changes what the system detects, for everyone, and it does so silently. The lower-level events are no longer alerts at all.

The three techniques above leave detection alone. The data still says the error rate rose. They change how people are told: once instead of many times, as one incident instead of many alerts, or not at all until something else changes. That keeps them reviewable and reversible. You can look at what was merged, grouped or held back, and you can change the rule tomorrow. Raise a threshold only for an alert that needs action but fires too early. Use the other tools for noise.

The order to apply them

The order runs from the safest to the riskiest.

  1. Deduplicate first. When the key is right it loses nothing, and it removes the most obvious repetition.
  2. Correlate second. It reduces many alerts to few incidents. It is a judgement, so keep it inspectable.
  3. Suppress last. By then the volume is smaller and each remaining alert is easier to judge. Each rule needs a reason and a review.

An illustrative example, with invented numbers, shows the effect on one bad hour:

StageAlerts or incidentsWhat changed
Raw notifications200Every check result that failed
After deduplication60Repeats merged into one alert each, with a count
After correlation8Related alerts grouped under a shared cause
After suppression3Known consequences and a flaky alert held back

Three pages is a workable number for a person. It is also three pages the team can explain, with each dropped step recorded.

Choosing among them

Ask what is wrong. If the same thing repeats, deduplicate. If different alerts share a cause, correlate. If an alert is real but should not reach a person now, suppress it by a rule with a reason. If an alert fires when there is nothing to find, change or remove the alert. And if the pattern is a general flood, take it to the wider alert fatigue plan, because tools reduce noise and only the team’s rules keep it down. Decide, too, where those rules end up: a person who is paged should still be able to see what was held back and why. The escalation policy is the place that decides who is told.

What to look for in a tool

Whatever software does this work, a few questions separate a tool you can trust from one you have to take on faith.

  • Can you see the key and the grouping rule? A dedup key or a grouping reason you cannot read is a black box.
  • Is nothing thrown away? Merged, grouped and suppressed alerts should stay in the data, with counts and timestamps.
  • Can a person override it? Splitting an incident and un-suppressing an alert should each take one step.
  • Do rules expire? A suppression with no end date is a decision nobody will revisit.
  • What are its limits? Ask how it behaves when the data is missing or the topology is out of date, because that is when grouping quietly goes wrong.

Frequently asked questions

What is the difference between deduplication and correlation?

Deduplication treats alerts that are the same as one alert, usually because they share a key such as the source, the check and the target. Correlation treats alerts that are different but related as one incident, because they probably come from the same cause. Deduplication is exact; correlation is a judgement.

Is suppression the same as muting?

Muting is one kind of suppression, usually manual and temporary. Suppression in general is any rule that stops a notification: a maintenance window, a low-priority rule at night, or a rule that holds back an alert while its upstream cause is already firing. The defining feature is that the alert exists but nobody is told.

Why not just raise the threshold?

A threshold decides whether the condition counts as an alert at all. Raising it changes what you detect, for everyone, for good. Deduplication, correlation and suppression change what people are told while keeping the alert on record. That makes them easier to review and to reverse.

Can correlation hide a second problem?

Yes. If grouping is too loose, an independent failure that happens at the same time is folded into the incident already open, and nobody looks for it. Keep the grouping rule specific, let a person split an incident, and watch for alerts that arrive after the cause is fixed.

Should machine learning do the correlation?

It can help, but it has limits. A model can find patterns across many alerts that a fixed rule would miss, and it can also group things that only look alike. Prefer grouping that can explain itself, and keep a way to see and correct what was grouped.