MTTA vs MTTR

Published Updated 7 min read

MTTA, mean time to acknowledge, is the average time from an alert paging someone to a person acknowledging it. MTTR usually means mean time to resolve: the average time from the start of an incident to its resolution. MTTA measures how fast people respond; MTTR measures how fast the problem goes away.

The two definitions

MTTA is the sum of the time between each page and its acknowledgement, divided by the number of pages. It starts when the page goes out, not when the problem began, and it ends when a person says “mine”.

MTTR is the sum of the time from each incident’s start to its resolution, divided by the number of incidents. The R is ambiguous: teams use it for repair, recovery, restore, resolve and even respond. Write down which one you mean, and where the clock starts: when the problem began, when it was detected, or when the incident was opened. Those can be very different.

Where each clock starts and stops

An incident has a sequence of moments, and each metric is a slice of it.

MetricThe clock runs from, toWhat it measures
MTTD, time to detectThe problem begins, to an alert firing or a person noticingYour monitoring
MTTA, time to acknowledgeThe page going out, to a person acknowledging itThe paging path and the responders
MTTR, time to resolveThe incident starting (or being opened), to the problem goneThe whole response

The slices overlap on purpose. A slow MTTR with a fast MTTA points at diagnosis and the fix. A slow MTTA points at who is paged, how, and how many pointless pages came first. A slow MTTD is a monitoring problem that no response speed can repair. Finding what nobody watches is the subject of MTTD and blind spots.

How to measure them honestly

  • Use the median as well as the mean. One six-hour outage drags an average for months. The median shows the typical incident; the mean shows the pain.
  • Split by severity. A fast MTTR across hundreds of small alerts can hide slow recoveries from the incidents that matter. Report each severity on its own.
  • Measure from timestamps, not memory. The page, the acknowledgement and the resolution should all be recorded by the tools involved, not typed into a review afterwards.
  • Watch for acknowledging to stop the noise. If people acknowledge pages just to silence them, MTTA looks great while nobody is working the problem. Look at time to first action as well.
  • Keep the definition fixed. Changing where the clock starts changes the number. If you must change it, restate the earlier months.

Why the mean misleads

This example is invented. Suppose three incidents last 4, 6 and 170 minutes. The mean is 60 minutes, a figure that describes none of them. The median is 6 minutes, which describes two of the three and says nothing of the long one. Neither number is wrong. Each answers a different question, which is why a dashboard with one figure is not enough.

Durations of incidents are usually lopsided in this way: many short ones and a few very long ones. So the mean moves when a single bad incident happens, and it can improve for reasons unrelated to any change you made. Treat a change in MTTR as a prompt to look at the incidents behind it, not as proof of improvement or decline.

What moves each one

MTTA moves with the paging path: whether the alert reaches the right person, through a channel that wakes them, with short enough timeouts before the next person is paged. That is the work of an escalation policy. Noise matters too: people who are paged for flaky alerts all week answer more slowly when the real one comes, and a 30-day plan for alert fatigue is the way to cut it.

MTTR moves with diagnosis. Often, the fix is quick once the cause is known, and the time goes into finding it. Anything that puts the relevant metrics, logs and recent changes in front of the responder in the first minutes, such as good runbooks, clear service ownership or an AI SRE that investigates alongside paging, shortens it.

What the numbers tell the business

Customers feel MTTR, not MTTA. The minutes a checkout is down or an API returns errors are what show up in support tickets, refunds and renewal conversations, so MTTR by severity is the reliability number worth putting in front of the business.

MTTA is an early signal about the team. A rising MTTA can point to tired people, noisy alerts or a rotation that is too thin, and it can move before retention does.

Both are averages over very different incidents. When they go to leadership, pair them with the count and the customer impact of the worst incidents in the period, so one quiet month does not hide one expensive outage.

Traps that make the numbers lie

  • Alerts that resolve themselves. A flaky alert that clears in a minute is not an incident. Counted as one, it pulls MTTR down and hides the real ones. Separate self-resolving alerts from incidents before you average anything.
  • Closing early. If a team is judged on MTTR, an incident can be closed the moment the graph looks better and reopened later as a new one. Decide when an incident is resolved, and hold to it.
  • Splitting and merging. One outage that raises ten alerts is one incident. Counting it as ten distorts every average; counting ten separate problems as one hides them. Group alerts into incidents on purpose.
  • Comparing with other teams. Definitions differ, and so do systems. Compare a team with its own past, not with a number from somewhere else.

How to start measuring

  1. Write the definitions. Which R you mean, and where each clock starts and stops.
  2. Record the timestamps in the tools. The page, the acknowledgement, the first action and the resolution.
  3. Report by severity, with the median and the mean. Add the count of incidents, so a small sample is visible.
  4. Review the worst incidents each month. The numbers point at where to look; the incidents explain what happened.

Ten minutes on this each month is more useful than a dashboard nobody reads.

Which one to watch

Both, for different reasons. MTTA tells you whether your on-call setup works. MTTR tells you whether your systems and your team recover well. A team with a good MTTA and a poor MTTR is responding fast and then struggling to understand what it is looking at. A team with a poor MTTA and a good MTTR has capable engineers who take too long to be reached.

Track them by severity, keep the definitions fixed, and read them next to the incidents behind them. Compensation and rotation health feed both numbers, so the on-call compensation question belongs in the same conversation.

Frequently asked questions

What does the R in MTTR stand for?

It has been used for repair, recovery, restore, resolve and respond, and they are not the same interval. Pick one, write it down, and say where the clock starts. A team that mixes definitions cannot compare its own months, let alone another team's numbers.

Is a lower MTTA always better?

Not by itself. A low MTTA can mean people are responding quickly, or that they acknowledge pages just to silence them. Read it beside the time to first action, and check that acknowledged incidents are actually being worked.

Should we report the mean or the median?

Both. Incident durations are lopsided: a few long incidents pull the mean far above the typical case. The median shows what a normal incident looks like, and the mean shows how much the worst ones hurt. Reporting one without the other hides half the picture.

How do we lower MTTA?

Fix the paging path first: reach the right person through a channel that wakes them, with timeouts short enough that the next person is paged quickly when nobody answers. Then cut noise, because people who are paged for flaky alerts all week answer the real one more slowly.

What is the difference between MTTD and MTTA?

MTTD, mean time to detect, runs from the start of a problem until an alert or a person notices it. MTTA starts later, when the alert pages someone, and ends when a person acknowledges. MTTD measures monitoring; MTTA measures the paging path and the people behind it.