Learn: on-call, alerts and AI in SRE

Plain guides for the people who carry the pager and the people who answer for it.

Subscribe by RSS

AI in SRE

What AI agents do during an incident, where a human stays in the loop, and how to judge one before it goes near on-call.

  • What is an AI SRE?

    An AI SRE is software that does the first-response work of a site reliability engineer during an incident. It reads the alert, checks telemetry, logs and recent changes, forms a hypothesis about the cause, and reports it to the people on call. A good one shows its evidence and asks before it changes anything.

    Matan Cohen 8 min read

  • Human-in-the-loop for incident agents: where the approval gate goes

    Put the approval gate where an agent's action is costly or hard to undo. Let an incident agent read freely, ask before it makes a reversible change, and never take an irreversible step alone. Decide each gate by blast radius and reversibility, name who approves, and settle in advance what happens when nobody answers.

    Genady Tsvik 7 min read

  • Read-only first: a permissions model for production AI agents

    Start a production AI agent with permission to read and nothing else, then grant writes one narrow scope at a time. Give the agent its own identity, keep each credential to the least it needs, log every call, and make revoking access one quick step. What an agent cannot reach, no instruction can make it change.

    Matan Cohen 7 min read

  • How to evaluate an incident agent before it touches on-call

    Evaluate an incident agent by replaying past incidents it has not seen and scoring its findings against the cause your team later established. Judge whether its evidence is right, not how confident it sounds. Then move it from shadow mode towards live on-call in stages, with the exit criteria written down before you start.

    Genady Tsvik 7 min read

  • The ironies of automation, applied to on-call AI

    The ironies of automation are a set of problems named in a 1983 paper: the more of a job a machine does, the harder the leftover human job becomes. For on-call AI, that means skills fade, watching an agent is harder than doing the work, and a failure hands the engineer the hardest incident with the least practice.

    Matan Cohen 7 min read

  • AI-written postmortems: draft yes, lessons no

    An AI can draft the factual part of a postmortem: the timeline, the changes involved and the impact, assembled from the record. It cannot supply the lessons. Deciding why people acted as they did, what to change and who owns it takes human judgement, in a blameless review that the team holds together.

    Genady Tsvik 7 min read

  • An AI voice agent that phones the on-call engineer: how it works

    A voice agent for on-call phones the on-call engineer on its own, as the escalation policy or workflow defines, and briefs them on the incident. It suits urgent problems, because a call is harder to miss than a message. Its limits are audio only, speech recognition, and the need for checks before consequential actions.

    Matan Cohen 7 min read

Noise and alerts

Why alerts turn to noise, what deduplication, correlation and suppression each do, and how to cut the noise with rules you can defend.

  • Alert fatigue: a 30-day plan

    Alert fatigue is what happens when people get more alerts than they can act on, so they stop trusting any of them. A 30-day plan fixes it in order: audit what fires, sort each alert as actionable or informational, delete or reroute the rest, tune what remains, and review every week. Give it an owner, or nothing changes.

    Genady Tsvik 8 min read

  • Deduplication, correlation and suppression: what is the difference?

    Deduplication merges identical repeated alerts into one. Correlation groups different alerts that share a cause into one incident. Suppression deliberately withholds a notification by a rule. Deduplication and correlation change how alerts are presented. Suppression withholds a person's attention, so apply them in that order, and never confuse any of them with raising a threshold.

    Matan Cohen 7 min read

  • Suppress by policy: alert rules a human can defend

    Suppress alerts by policy when a written rule, not a habit, decides that nobody is paged. A defensible rule says what it matches, why holding it back is safe, who owns it and when it ends, and it leaves the alert visible for review. If you cannot explain a rule to a colleague at 3 a.m., do not suppress.

    Matan Cohen 8 min read

On-call and escalation

Escalation policies, the channels that wake people, and running a fair rotation with a small team.

  • What is an escalation policy?

    An escalation policy is the set of rules that decides who gets paged for an alert, in what order, and how long each responder has to acknowledge before the next person is paged. A good policy always ends with someone who will answer, so an alert is never left without a responder.

    Genady Tsvik 7 min read

  • Which channel wakes an engineer: phone, SMS, push or Slack

    A phone call is the channel most likely to wake a sleeping engineer, SMS is the most widely reachable, push depends on an app that is installed and signed in, and chat is for people who are awake. Each fails in its own way, so layer them, require an acknowledgement, and test the whole path from the phone.

    Matan Cohen 7 min read

  • Escalation policy anti-patterns, and how to test yours

    The common escalation policy anti-patterns are a single point of failure, timeouts that are too long or too short, paging everyone at once, stale schedules, and no backup for the backup. Find them by testing: page the policy on purpose, let each step time out, and read the schedule behind it.

    Genady Tsvik 7 min read

  • On-call for a team of five

    Five is a floor for a sustainable rotation: one primary each week and a secondary as backup, so each person carries the pager about one week in five. It works only while page volume is low, handoffs are clear and leave is planned. When pages grow, fix the noise first, then hire or buy cover.

    Miki Manor 7 min read

  • On-call burnout: the signs and the fixes

    On-call burnout shows up as slower acknowledgements, muted or ignored alerts, swap and opt-out requests, and people leaving. Leaders can see these signals in the data and in how the rota behaves. Fix it by cause: cut noise, lower frequency, make the schedule predictable, and give people time to recover.

    Miki Manor 7 min read

  • On-call compensation

    On-call compensation is how a company pays people for being available to respond outside working hours. The common models are a flat stipend per shift or week, pay for each hour on standby, pay or time off for each page answered out of hours, or a mix. In some countries, the law decides part of it.

    Miki Manor 7 min read

Incident response and reliability

The numbers worth tracking, who runs an incident, and how to keep detecting what you do not monitor.

  • MTTA vs MTTR

    MTTA, mean time to acknowledge, is the average time from an alert paging someone to a person acknowledging it. MTTR usually means mean time to resolve: the average time from the start of an incident to its resolution. MTTA measures how fast people respond; MTTR measures how fast the problem goes away.

    Miki Manor 7 min read

  • The incident commander in a small company

    An incident commander is the one person who coordinates an incident: who is doing what, what is known, and what the outside world is told. In a small company the role can rotate among a few senior engineers, with the commander kept apart from the person fixing the problem. Declare early, hand over cleanly, and keep the role short.

    Miki Manor 7 min read

  • Running an incident channel in Slack or Teams

    Give each incident its own Slack or Teams channel, opened at declaration and named so it can be found later. Put the roles and a pinned summary at the top, keep the channel for facts and decisions, move side talk to threads, update on a set rhythm, hand over in the open, and close the channel with a final summary.

    Genady Tsvik 7 min read

  • MTTD and blind spots: why detection time matters

    MTTD, mean time to detect, is the average time from a problem starting to someone or something noticing it. It matters because nothing else can begin until detection does. Measure it from the real start, note who noticed first, and review regularly what you do not watch, since the slowest detection is the problem nobody is looking at.

    Miki Manor 7 min read

Switching

Moving your on-call to a new tool without a missed page.

  • Moving your on-call to a new tool without a missed page

    To move on-call to a new tool without a missed page, inventory what pages whom today, map schedules, policies and alert sources, run both tools side by side, and test with drills. Then cut over one team at a time with a rollback ready, and keep the old tool until the new one has proved itself.

    Miki Manor 8 min read