The ironies of automation, applied to on-call AI
The ironies of automation are a set of problems named in a 1983 paper: the more of a job a machine does, the harder the leftover human job becomes. For on-call AI, that means skills fade, watching an agent is harder than doing the work, and a failure hands the engineer the hardest incident with the least practice.
What Bainbridge described in 1983
Lisanne Bainbridge’s paper Ironies of Automation, published in the journal Automatica in 1983, is about industrial process control, long before language models. Its point has lasted because it is about people. Four of its observations apply directly to on-call.
- The leftover job is the hard one. Designers automate what they can. What remains for the human is whatever nobody knew how to automate, which is often the difficult part.
- Skills fade with disuse. A person who no longer does the work regularly gets worse at it, yet is expected to take over when the automation fails.
- Monitoring is hard to do well. People are poor at staying alert to a system that is almost always right and rarely needs them.
- The more automation, the more the human matters. The situations that reach a person are the abnormal ones, so the quality of that person’s judgement counts for more, not less.
None of this argues against automation. It argues for designing around the person who is left.
How each irony appears in on-call
| Irony | What it looks like with an AI on-call agent | A countermeasure |
|---|---|---|
| The leftover job is the hard one | The agent clears routine incidents, so the ones people see are the strange ones | Prepare people for strange incidents, not routine ones |
| Skills fade with disuse | Engineers stop learning the systems by investigating them | Planned manual practice, and learning paths for new engineers |
| Monitoring is hard | People skim the agent’s output and approve it | Make findings quick to check; sample and review them |
| The human matters most at the handoff | The agent fails on a hard incident and passes it over with little context | Design and rehearse the handoff |
The table shows why “the agent handles the easy ones” is a weaker promise than it sounds. The easy incidents were where engineers built their picture of how the system behaves.
Skills fade, and juniors never form
The first-response work of an incident, reading the alert, finding the dashboard, tracing the request, is also how engineers come to understand a system. If an agent does that work, the understanding does not form by accident any more. Experienced engineers keep what they have for a while. New engineers may never build it.
The response is to treat understanding as something the team has to produce on purpose:
- Have new engineers investigate past incidents by hand, then compare with what the agent found.
- Run regular practice sessions with the agent switched off, using real past incidents.
- Rotate who reviews the agent’s findings in depth, so the skill stays spread across the team.
- Keep runbooks and service maps current, because they are what a person falls back on.
Watching an agent is harder than doing the job
Reviewing an agent’s investigation sounds easier than performing it. Often it is harder. To judge a conclusion you need to know what evidence would support it and what evidence is missing, which takes about the same understanding as the investigation itself, plus the effort of reading someone else’s reasoning. When the agent is usually right, that effort drops, and review turns into approval. This tendency is called automation bias.
Design for the reviewer:
- Show the path, not only the answer. The queries run, the data returned, the alternatives considered.
- Make each claim checkable in one step. A link to the log line or the change, not a paragraph of summary.
- State what was not checked. An honest gap is a prompt to look there.
- Sample and score. Regularly check a set of findings against what was later found, as described in how to evaluate an incident agent.
Why a language model sharpens the problem
Classic automation can fail quietly too, by masking a developing fault until it hands over. A language model adds a new kind: its output is equally fluent whether it is right or wrong, the same alert can produce different answers on different days, and a change in the model, the prompt or the data behind it can shift quality without any error being raised. Confidence in the wording is not a signal of accuracy.
That makes the monitoring irony stronger. The reviewer cannot rely on the agent to warn them when it is out of its depth, so the team has to build that check from outside: regular scoring against known outcomes, visible limits on what the agent can read, and a habit of asking what the tools could not see.
The handoff when the agent fails
An agent is most likely to fail on the incidents that are unusual, which are also the ones that matter. When it does, the engineer has to take over cold, with the incident already underway. Three things make that better or worse.
- How early it says so. An agent that reports “no cause found” after ten minutes is more useful than one that circles for an hour. Agree a limit on time and effort, after which it hands over.
- What it hands over. A running summary: what it checked, what it found, what it ruled out, what it did not check, and what it would look at next. The engineer should not have to reconstruct the session.
- What state it leaves. Any change it made or proposed is listed. The gates in human-in-the-loop for incident agents keep that list short.
Also monitor the agent itself. If it is down or degraded, on-call engineers need to know they are on their own, before an incident, not during one.
Keeping people ready
Some practical rules follow from the ironies, all of them about the team and its tools:
- Practise the manual path. A drill with the agent off costs an afternoon and shows what the team can still do.
- Give people real work. Investigation that only ever confirms the agent’s answer teaches nothing. Let engineers form their own hypothesis first sometimes.
- Measure people as well as the agent. Track how long the team takes without the agent, not only with it.
- Keep the model’s limits visible. Everyone should know what data the agent can reach and where it is blind. The first part of what an AI SRE is lists what it needs to work.
A team that does these things gets the benefit of the agent and keeps the skill to manage without it.
Frequently asked questions
What are the ironies of automation?
They are a set of contradictions described by the researcher Lisanne Bainbridge in a 1983 paper on automated industrial control. Automating a task tends to leave people the parts nobody could automate, weakens the skills they need for those parts, and asks them to monitor a system that rarely gives them anything to do.
Does this mean we should not use AI in on-call?
No. It means design for the human who is left. The ironies describe what goes wrong when automation is added without thinking about the person who must take over. With deliberate practice, good handoffs and agents that show their work, the benefits can be kept and the risks reduced.
What is automation bias?
It is the tendency to accept an automated system's output without checking it, especially when the system is usually right. On-call AI invites it: a fluent, plausible finding at 3 a.m. is easy to trust. The countermeasure is to make checking quick and normal, so that trusting is a decision.
How do junior engineers learn if the agent does the first investigation?
Deliberately. Have new engineers investigate past incidents by hand before they see the agent's work, pair them with experienced responders, and run practice sessions with the agent switched off. Experience that used to arrive by accident now has to be planned.
What should the agent do when it cannot handle an incident?
Say so early, plainly, and hand over with a summary: what it checked, what it found, what it did not check and what it suggests next. The handoff is the moment the ironies bite, so design it and test it before you need it.