AI-written postmortems: draft yes, lessons no
An AI can draft the factual part of a postmortem: the timeline, the changes involved and the impact, assembled from the record. It cannot supply the lessons. Deciding why people acted as they did, what to change and who owns it takes human judgement, in a blameless review that the team holds together.
What a postmortem is for
A postmortem is a way to change what happens next time. The document is a by-product. The point is that a team looks honestly at an incident, agrees what made it possible, and commits to changes that make the next one less likely or less painful. The practice of doing this without blame is set out in the postmortem culture chapter of the Google SRE book: people are assumed to have acted sensibly on what they knew, and the review looks at the system around them.
This matters for AI because writing is not the hard part. The hard part is the conversation that produces the content. Any tool that makes the document appear without the conversation solves the wrong problem.
What a model can do, and what it cannot
A language model is good at assembling and summarising a record. It is poor at judgement about why people did what they did. The split follows from that.
| Part of the postmortem | A model can draft it | A person has to decide |
|---|---|---|
| Timeline | Yes: from alerts, chat, deploys and the incident record | Whether it is complete and the times are right |
| Impact | A first summary from metrics and tickets | What the impact meant to customers and the business |
| What was changed and by whom | A list from deploy and change history | Which changes mattered |
| The narrative | A plain account of what happened | Whether it is fair to the people involved |
| Contributing factors | Suggestions to react to | What actually contributed, including process and habit |
| Actions | Candidates, drawn from the discussion | Which to do, who owns them, by when |
| What was learned | Nothing reliable | Everything |
The last row is the reason for the title. A model can draft; it cannot learn on the team’s behalf.
An example of the split
This example is invented. A bad configuration push goes out on a Tuesday and the customer portal returns errors for a while until it is rolled back.
The model’s draft states that the change was deployed at a given time, that errors began soon after, that the rollback completed later, and which alerts fired. Each line links to a deploy record, a graph or a message. The engineer who led the incident corrects one time and removes a guess about the duration.
The team meeting then adds what no record contained. The change had passed review because the reviewer could not see the production values. The rollback took longer than it should have because the runbook was out of date. Two people had noticed the errors early and each assumed the other had raised them. Those three observations lead to three actions: show real values in review, fix the runbook, and agree who declares an incident. None of them was in the draft, and all of them are the reason the review was worth holding.
The risks of a fluent draft
A draft that reads well is easy to accept. Four risks follow from that.
- Invented detail. A model can produce a time, a service or a quote that fits the story and is not in the record. Every fact needs a source.
- One tidy cause. Models like a clean narrative. Real incidents usually have several contributing factors, and a single “root cause” can close the discussion too early.
- Hindsight. A draft written after the fact can make the right action look obvious. The review has to reconstruct what was known at the time.
- Blame by wording. A sentence like “the engineer failed to check” assigns fault. Ask the draft to describe actions and circumstances, not people’s failures.
The design of the draft can reduce these. A good one cites its sources for each statement, keeps facts apart from interpretation, marks uncertain parts, and avoids naming individuals as causes.
A workflow that keeps the learning
The rule is simple: the model prepares, people conclude. A workflow that follows it:
- Draft the record. From the incident channel, alerts, deploys and the timeline, the model prepares a draft of the timeline, the impact and the changes involved, marked as unverified.
- Verify the facts. The person who led the incident checks each fact against its source and fixes what is wrong.
- Meet before concluding. The team discusses what happened and why, working from the verified facts. Ask what made each decision sensible at the time and what information was missing.
- Write the analysis by hand. People write the contributing factors and the lessons, in their own words. Then compare with any suggestions the model produced: agreement is reassuring, and disagreement is worth a look.
- Assign the actions. Each one gets an owner and a date, and a place where it is tracked.
- Close the loop. At the next review, check that earlier actions were completed, and whether they helped.
Blameless in practice
Blameless does not mean without accountability. It means the review asks better questions. Some that work:
- What did the person know at that moment, and why did the action make sense?
- What made the right action hard to see or hard to take?
- What was missing: a dashboard, a runbook, a permission, a colleague awake?
- What would have made this incident smaller if it had happened tomorrow?
These are questions for people in a room. A model can help afterwards: searching earlier reviews for repeated themes, or checking that an action list matches the discussion. That is a useful supporting role, and it keeps the judgement with the team.
Before you adopt it
A short set of rules keeps AI drafting from eroding the practice:
- A named owner for every review, who signs off the facts.
- A template the team designed, so the draft has the right shape and the analysis stays human.
- Provenance. Each statement in the draft links to its source, or is labelled as a summary.
- Review of the drafts. Treat the model like any other tool that writes: sample its output and score it, as in how to evaluate an incident agent.
- Low-stakes access. Drafting a document is a low-risk write, and it should run within the gates described in human-in-the-loop for incident agents.
A team that keeps these rules can save real time on the record and still spend its effort on the part that changes anything.
Frequently asked questions
Can an AI write a whole postmortem?
It can write a full draft, but the draft is not the postmortem. The facts can be assembled from the record and checked. The analysis, the contributing factors, the actions and the owners have to come from the people who were involved, because the value of the exercise is that the team reaches them together.
What is a blameless postmortem?
A review that looks at how a system, a process and a set of circumstances led to an incident, not at which person is at fault. The premise is that people acted sensibly with the information they had. Blame teaches people to hide problems, and blameless review keeps them visible.
What can go wrong when a model writes the draft?
It can invent a detail, such as a time or a service name, that reads as true. It can settle on one tidy cause when several factors combined. And it can make the review feel finished before the team has met. Verify every fact against its source, and discuss before you conclude.
How should we label AI-written parts?
Mark them as drafts until a named person has checked them, and say what they were drawn from. Readers should be able to tell which statements were verified against logs and chat, and which are the model's summary of them.
How does the team keep learning if a model does the writing?
Keep the meeting, the discussion and the choice of actions with people, and use the model for the record and for searching earlier reviews. Track whether actions are completed. A postmortem that produces a document but no change has not done its job, whoever typed it.