An AI voice agent that phones the on-call engineer: how it works

Published Updated 7 min read

A voice agent for on-call phones the on-call engineer on its own, as the escalation policy or workflow defines, and briefs them on the incident. It suits urgent problems, because a call is harder to miss than a message. Its limits are audio only, speech recognition, and the need for checks before consequential actions.

What happens on the call

Voice paging replaces one step of an escalation policy: instead of a text or a push notification, a phone call. The steps of a well-built one are these.

  1. An alert is routed to a policy. The escalation policy chooses who is on call and how to reach them.
  2. The call is placed. The voice agent calls the on-call engineer on its own, as the escalation policy or workflow defines. Nobody has to start it.
  3. It identifies itself and briefs. It says who is calling, what is wrong, who is affected, since when, and what has already been found.
  4. The engineer responds. They can ask for detail, or tell it what to do next.
  5. The outcome is recorded. The transcript, the result of the call and a summary go on the incident timeline, so the team can see what was said and decided.

If the engineer does not answer, the next step of the escalation policy pages the next person. That decision belongs to the policy, which the team wrote, and not to the agent.

How it works under the hood

The architecture is a chain of parts, each with its own limits. A telephony provider places the call and carries the audio. Speech recognition turns what the engineer says into text. A language model decides what to say or do next, working with tools that read live incident data, such as the alert, the service, recent changes and what other agents have found. Speech synthesis turns the reply back into audio. Around all of it sits state: which incident, which step of the policy, what has been said.

Two consequences follow. First, latency adds up across the chain, so a call feels natural only when each part is quick. Second, an error in any part reaches the engineer as a mistake in the conversation, whether it is a misheard word, a wrong fact from a tool or a slow reply. The tools and the data behind the briefing matter more than the voice.

When a phone call is the right channel

A call is among the hardest channels to ignore, and that has a cost: it interrupts. The other channels and how each fails are compared in which channel wakes an engineer. Use a call where the interruption is worth it.

SituationCall?Why
Urgent, customer-facing failure at nightYesThe hardest channel to sleep through
An earlier, gentler step went unansweredYesThe policy has already tried the quiet options
A decision is needed, and talking is fasterYesThe engineer can ask questions and get answers in seconds
A warning that can wait until morningNoA message in a channel is enough
A flaky alert that resolves itselfNoCalls for noise teach people to ignore calls

Calling for everything wears the channel out. The rest of the policy, which alerts are urgent and which are not, decides whether a call still means something when it comes.

The briefing

Audio is linear. The engineer cannot skim, scroll back or open a graph, so a briefing has to be short and in the right order. A useful order is: what is broken, who is affected, since when, what has been found, what is suggested, and what is needed from the engineer.

This is an illustrative example of the shape, not a script:

This is the incident agent with an urgent alert for the customer sign-in service. Sign-ins have been failing for six minutes. The errors began when a certificate expired on the gateway. The suggested action is to renew it, which needs your approval. Do you want more detail, or should the platform team be paged?

It names the service, gives the impact and the time, states the finding, and ends on a choice the engineer can answer in a word. It repeats nothing it does not need to. A written version, with the links and the evidence, should go to the incident channel at the same time, because a call is a poor place for detail.

Checking before consequential actions

Some things the voice agent can do during a call are trivial, and some are not. The voice agent checks with the engineer before paging someone else, escalating, downgrading or muting a monitor. Each of those affects other people or hides a signal, so a person should decide, and the agent should say back what it is about to do before it does it.

A spoken yes needs care. Speech recognition can mishear, and a person half awake can agree to something they did not follow. Good design reads the action back in plain words, asks for a clear confirmation, and keeps the action list short. How to decide which actions need a person is covered in human-in-the-loop for incident agents.

Limits to plan for

  • Audio only. No graphs, no long identifiers, no tables. Anything the engineer must read goes to text.
  • Recognition errors. Names, numbers, service names and accents can all cause mistakes. Repeat key facts and confirm anything that matters.
  • The line. Phones can be off, silenced, out of signal or set to do-not-disturb. Ask each engineer to allow calls from the number the agent uses.
  • Speed. A reply that takes a second or two too long makes a call feel broken.
  • Recording rules. Check what your jurisdiction requires before you record or transcribe.
  • No substitute for a plan. A call gets an engineer to the incident. What happens next depends on the runbooks, tools and access the engineer has.

Questions to ask about any voice agent

  • Who decides when it calls, and where is that written down?
  • What does it say, and how long does it take to say it?
  • What can it do during the call, and which actions does it check first?
  • Where does the record of the call go?
  • What does the engineer see in text at the same time?
  • What happens when nobody answers?

For a first look at the whole agent behind the voice, read what an AI SRE is.

Frequently asked questions

How does an AI voice agent know when to call?

It does not decide that on a whim. The call is a step in an escalation policy or a workflow, which your team defines: which alerts, which people, in which order. The agent carries out that step and then does the part a person would otherwise do, telling the engineer what is happening.

Why use a phone call instead of a message?

A call is harder to miss than a message, and it can get an answer in the moment. It is not a guarantee: phones silence or filter calls as well as notifications, and a call can reach voicemail. For urgent incidents, especially at night, a call is the channel that is hardest to sleep through. For low urgency, a message is kinder.

Can the voice agent make changes during the call?

It can take actions, within the access it has been given, and it should check with the engineer before paging someone else, escalating, downgrading or muting a monitor. Those actions affect other people or hide a signal, so a human decides. Everything it does should land in the incident record.

What can go wrong on a voice call?

Speech recognition can mishear a name, a number or a service. The line can be noisy or drop. The engineer cannot see a graph. And a call is one long stream of audio with nothing to skim. Good design keeps briefings short, repeats key facts, and sends the detail as text as well.

Is it legal to record and transcribe on-call calls?

The rules on recording and transcribing calls differ by country and, in some places, by state, and they can require consent from everyone on the call. Check with a lawyer where your engineers are based, and tell people plainly that the call is recorded or transcribed. This page is not legal advice.