MTTD and blind spots: why detection time matters

Published Updated 7 min read

MTTD, mean time to detect, is the average time from a problem starting to someone or something noticing it. It matters because nothing else can begin until detection does. Measure it from the real start, note who noticed first, and review regularly what you do not watch, since the slowest detection is the problem nobody is looking at.

What MTTD is

Every incident has a beginning that nobody saw. A deploy goes out, a certificate expires, a queue starts to back up. MTTD, mean time to detect, is the average gap between that beginning and the moment someone or something notices. Until then, the incident is happening and nobody is responding to it.

Detection is not the same as an alert. An alert is one route to detection. A colleague noticing that the dashboard looks strange is another. A customer writing to support is a third, and the least welcome. The measure is about the first moment anyone knew, by any route.

Why detection time matters to the business

Everything that comes after detection waits for it. The page cannot go out, the acknowledgement cannot happen and the fix cannot begin. MTTA and MTTR are measured from moments that only exist once someone has detected the problem, so a team can respond quickly and still be late overall.

The cost is easiest to see through customers. When a problem is found by monitoring, the team is fixing it while customers may notice little. When it is found by customers, they have already been affected, some have written to support, and some have decided what to think of the company. Late detection turns into refunds, credits, lost renewals and a reputation for being the last to know. For the business, the cost of a minute of undetected failure is paid in trust, and it grows with each customer who sees it first.

Measuring it honestly

MTTD is harder to measure than it looks, because the start of a problem is not recorded anywhere at the time. It has to be reconstructed, usually in the review, from the first sign in the data.

  • Record two times. The first anomaly in logs or metrics, and the first customer-visible impact. Say which one you call the start, and keep to it.
  • Note who noticed first. This is the most useful field, and easy to record.
  • Use the median and split by severity, for the same reasons as with response times, set out in how to measure MTTA and MTTR honestly.

The “who noticed first” record turns into a simple table:

Noticed first byWhat it tells you
An automated alertMonitoring covers this path
An engineer or colleagueLuck, or an informal check, not a reliable route
Support, from a ticketThe gap is in monitoring or in how tickets are triaged
A customer, publiclyThe most costly route, and a clear blind spot
A partner or third partyThe problem crossed a boundary you do not watch

Count how many incidents fall in each row. The trend in the lower rows matters more than the average time.

Detected is not the same as alerted

An alert that fires and is ignored has not detected anything. If people have learned to skim alerts because most are noise, the real one can sit in a channel for an hour. The remedy is the plan described in the 30-day alert fatigue plan: fewer, better alerts that people trust. Detection needs both the signal and the attention, and a team can lose either.

Where blind spots hide

A blind spot is a part of the system, or of the customer experience, that nobody watches. They cluster in predictable places.

  • New services. Launched without dashboards, alerts or an owner.
  • Third-party dependencies. A payment or email provider you rely on, watched by nobody on your side.
  • Failures by absence. A nightly job that did not run, or data that did not arrive, produces no error to alert on.
  • Rare journeys. Billing, password reset, exports and other paths customers use seldom but care about.
  • Partial degradation. Slow rather than down, for some users, some regions or some devices.

The dangerous common feature is silence. A blind spot does not show up as a red graph. It shows up as a customer message.

Keeping on finding them

Blind spots are found by looking, and by asking customers and past incidents where to look. Five habits do most of the work.

  1. A coverage map. List services and customer journeys down one side, the signals across the top: metrics, logs, a synthetic check, an alert, an owner. The empty cells are the gaps.
  2. Synthetic checks. Scripted journeys run from outside your network, such as “sign in and pay”, that fail when the customer’s experience fails. Add heartbeat checks for jobs that should run, so silence raises an alarm.
  3. Customers as sensors. Make sure a customer report reaches an engineer quickly, with support tagging tickets that look like outages.
  4. Ask of every incident: who noticed first? If it was not an alert, ask what check would have found it, and add it.
  5. A standing review. A short meeting each quarter, and a launch checklist that will not let a new service go live without a check and an owner.

An illustrative slice of a coverage map shows how the gaps read:

Service or journeyMetricsAlertSynthetic checkOwner
CheckoutYesYesYesPayments
Password resetYesNoNoIdentity
Nightly report exportNoNoNoNobody

The last row is the one to fix first: no signal, no alert and no owner.

An example: one late detection, followed through

This example is invented. A report export feature starts failing quietly on a Monday. Nothing alerts, because the export runs as a background job and has no check. On Thursday a customer asks support why their weekly report is missing. The team traces the failure to Monday, so detection took three days, and finds that the same job has no owner.

The review records the start time, who noticed first (a customer, through support), and why nothing fired. It produces three changes: a heartbeat check that raises an alarm when the export has not run, a named owner, and a line in the launch checklist for background jobs. The next time the job stops, the team knows within the hour. The fix did not need to be clever. It needed someone to ask who noticed first.

What MTTD hides

MTTD is an average over incidents that were found. It says nothing about the ones that were not. A team with a low MTTD can still have a system full of silent failures, because the failures nobody noticed never enter the calculation. That is why “who noticed first” and the coverage map matter more than the number.

Read it with the incident record: how many were found by customers, how long the slowest took, and what changed after the last review. When the incident commander and the team review an incident, add detection to the agenda next to response, since a fix for detection helps every future incident.

Questions for leaders

  • What would a customer notice before we did?
  • Which important journey has no check?
  • How many incidents last quarter were first reported by customers?
  • Does every service have a named owner and a way to be alerted?
  • What did we change after the last incident to find the next one sooner?

The answers are a shortlist for the next quarter, and they cost far less than a customer-visible outage that nobody saw coming.

Frequently asked questions

What does MTTD stand for?

Mean time to detect: the average time between a problem starting and it being noticed, by an alert, a person or a customer. It measures how quickly your monitoring and your people find problems, before any response begins. It is a different interval from MTTA, which starts when a page goes out.

How do we know when a problem really started?

Usually after the fact, from the first sign in the data: the first error in the logs, the first failed request, the first metric that moved. Record two times, the first anomaly and the first customer-visible impact, and be consistent about which one you call the start.

What is a monitoring blind spot?

A part of the system or of the customer experience that nobody watches, so a failure there goes unnoticed. Common ones are new services, third-party dependencies, background jobs that fail by not running, and rarely used journeys such as billing or password reset.

Who should own detection?

The team that owns each service, with a named owner for every service and every important customer journey. Detection that belongs to nobody is the usual source of blind spots. A short review each quarter keeps the owners and the checks current.

How often should we review for blind spots?

On a regular schedule, such as every quarter, and whenever a new service or a major feature launches. Add a standing question to each incident review: who noticed first, and what would have found it sooner? Those answers are the best source of new checks.