How to evaluate an incident agent before it touches on-call

Published Updated 7 min read

Evaluate an incident agent by replaying past incidents it has not seen and scoring its findings against the cause your team later established. Judge whether its evidence is right, not how confident it sounds. Then move it from shadow mode towards live on-call in stages, with the exit criteria written down before you start.

Why test before it is on-call

An incident agent will be wrong sometimes. What matters is how it is wrong, and whether a tired engineer at 3 a.m. can tell. A wrong answer stated calmly gets acted on. A wrong answer with a query, a log line and a note that says “not checked: the database” gets caught. You cannot learn which kind of agent you have from a demo, because a demo shows the cases that work. You learn it by testing on the cases that do not.

If you are new to the idea, what an AI SRE is sets out the work the agent is being asked to do. The test does not need special tools. It needs your own history, an honest score sheet and a team willing to read the results together.

Build the test set from your own incidents

Take past incidents and replay them. A useful set has a spread, not only the dramatic ones.

Case typeWhy it belongs in the setWhat to score
A clear cause, found quicklyThe agent should not miss the easy onesRight cause, right evidence
A cause found only after hoursShows whether it can reason past the obviousUseful direction, honest doubt
A flaky alert that resolved aloneTests triage: is this worth anyone’s time?Correct call, no invented drama
An incident with no known causeTests whether it admits it does not know“Unknown” stated plainly, checks listed
A change that looked guilty but was notTests whether it blames the obvious suspectWeighs the evidence, not the timing alone

For each case, record the ground truth from what the team learned: the cause, the fix, the timeline. Then replay the case with the agent seeing only what existed when the alert fired. If the agent can read the postmortem, or logs from after the fix, the score measures hindsight, not skill.

What to score

Score the parts of a finding separately, because an agent can be right for the wrong reason.

  • The hypothesis. Did it name the cause, a related cause, or a wrong one?
  • The evidence. Is each claim backed by a query, a log line or a change that exists and says what the agent says it says? Fabricated evidence is the most serious failure, and it is easy to find by checking.
  • What it left unchecked. Did it say what it did not look at?
  • The proposed actions. Were they safe, and in the right class for the gate they would face?
  • The usefulness. Would a responder who started from this finding have saved time, lost time, or been misled?

Use a plain rubric with a few grades: right, partly right, wrong, harmful. Two people should score each case, and any disagreement is worth a conversation, because it shows where the team’s own standards differ.

False confidence is the failure to watch

An agent that says “unknown, and here is what was checked” is doing its job. An agent that always has an answer is not useful, it is persuasive. Reward the first and penalise the second. Count the cases where the agent was wrong and sure, and treat each as a design problem to understand: was the data missing, the tool weak, the reasoning poor?

The same care applies to the people. When a team reads a finding, the wording, order and layout change how much they believe it. Part of the evaluation is checking that the way the agent presents its result invites checking, not deference.

Vanity metrics to leave out

Some numbers look good and prove nothing.

  • Investigations run, alerts handled, tokens used. Activity, not value.
  • Average confidence. Confidence is what you are testing, not evidence.
  • Time saved, estimated. Without a baseline from the same incidents, it is a guess.
  • A single accuracy figure. It hides the harmful cases inside the average.

Better measures come from the score sheet: how often the finding was right and well supported, how often the agent correctly said it did not know, how often a person changed or rejected what it proposed, and what an investigation cost.

Roll out in stages

Do not switch an agent from nothing to on-call. Move through stages, and give each one its own exit criteria.

StageWhat the agent doesWho sees itMoves on when
1. ReplayInvestigates past incidents offlineThe review teamThe score sheet meets the criteria
2. ShadowInvestigates live alerts; output kept out of the wayThe review team onlyIts findings match what people later establish
3. AdvisoryPosts findings to the incident; people decideOn-call engineersResponders rate it useful, evidence checks out
4. Approved actionsProposes changes that a person approvesOn-call, service ownersApprovals are given unchanged, nothing harmful
5. Narrow autonomyActs alone on a named, reversible patternEveryone, with an auditThe pattern has run cleanly and stays reviewed

Stages four and five depend on the controls in human-in-the-loop for incident agents and on the limits described in a permissions model. Test those controls too: try to make the agent exceed its access and see that it cannot.

Write the exit criteria first

Agree the criteria before the trial, when nobody has a favourite result. State what counts as passing, what counts as a reason to stop, and what happens then. Examples of a criterion: no fabricated evidence in any replayed case; every confident wrong answer reviewed and explained; the agent says “unknown” on the cases where the cause was never found; in the advisory stage, responders rate the findings useful and their evidence checks out, and in the approved-actions stage, requests are approved without edits and nothing harmful gets through. Add the rule for going back: if the record worsens, the agent steps back a stage until the team understands why.

Keep testing after launch

An agent changes when its model, its prompts or its tools change, and your systems change too. Rerun the replay set after each change, and add every new incident to it. The set becomes a regression test for the whole team’s confidence. Decide who owns it, when it runs, and who sees the result, the same way the team owns an on-call rota: a shared duty, a named owner and a regular review.

Frequently asked questions

Why not rely on a vendor demo or a benchmark?

A demo is chosen to work, and a benchmark measures someone else's incidents. Your systems, alerts and runbooks differ from both. The only test that predicts how an agent behaves on your on-call is one built from your own past incidents, scored by people who know what happened.

How many past incidents do we need?

Enough to cover the kinds of alert you actually get, including the awkward ones. A modest set with a clear range of cases teaches more than a large set of near-identical ones. Start small, score carefully, and add every new incident to the set as it closes.

What is hindsight leakage, and how do we avoid it?

It is the agent seeing information that did not exist when the incident began, such as the postmortem or later logs. Restrict each replay to the data available at the time of the alert, and keep postmortems out of anything the agent can read, or the score will flatter it.

What should we do with an agent that is often wrong but sounds sure?

Do not give it a role in on-call yet. Calm, confident errors are the most expensive kind, because people act on them. Look for whether it can say what it does not know, and whether every claim comes with evidence a person can check quickly.

When has an agent earned more responsibility?

When it has met the exit criteria of its current stage, written before the trial began, over enough cases and enough time to include the awkward ones. Widen its role one step at a time, keep the earlier controls, and be ready to step back a stage if the record worsens.