The incident commander in a small company

Published Updated 7 min read

An incident commander is the one person who coordinates an incident: who is doing what, what is known, and what the outside world is told. In a small company the role can rotate among a few senior engineers, with the commander kept apart from the person fixing the problem. Declare early, hand over cleanly, and keep the role short.

What the role is

During an incident, many things happen at once: people investigate, changes are tried, customers ask questions, and someone senior wants an update. The incident commander is the person who holds all of that together. They do not need to be the best engineer in the room. They need to know what is happening, who is doing it, and what has to happen next.

For a small company the role matters more than the size of the team suggests. With eight engineers, an incident can pull in everyone, and without a commander the result is a crowd: five people investigating the same log, nobody writing to customers, and no clear decision about whether to roll back. The commander turns a crowd into a team. A clear commander can shorten an outage, and customers judge the company on how long it lasts.

Who does it when there are eight engineers

A dedicated incident commander team is a luxury. A small company can do three things instead.

  1. Pick a handful of people. Choose three or four senior engineers, or people with good judgement under pressure, who share the role in turn. It does not have to be the most technical.
  2. Let the on-call engineer start. The first responder acts as commander for the first minutes: declares the incident, opens the channel, and names a commander. Then a designated person takes over, so the responder can go back to fixing.
  3. Avoid the obvious default. The founder or head of engineering is usually the person best placed to fix the problem, and their time is worth more there. They can be commander when nobody else is available, but they should not be the automatic choice.

Rotate the role on a schedule so that nobody carries it for long, and so that more people learn it. A role that only one person can fill is a single point of failure, in the same way as a pager with one name on it.

What the commander does, and does not

The list is short, and the second half matters as much as the first.

The commander doesThe commander does not
Declares the incident and sets its severityDebug or write code
Assigns roles and tasks to named peopleDo the task themselves because it is faster
Keeps a running picture of what is known and unknownArgue about the cause before the facts are in
Makes the decisions that need one ownerDecide alone what a specialist should decide
Makes sure stakeholders and customers are toldWrite every message; that is a separate role
Hands over when tired, and closes the incidentStay in the role for hours without relief

The habit to build is asking questions, not answering them. “Who is checking the database?” and “What do we know about the last deploy?” keep the incident moving without the commander touching a keyboard.

The other roles

On a small incident one person can wear several hats. On a larger one, separate the roles.

RoleJobCan be combined with
Incident commanderCoordinates and decidesNever the person fixing
Fixer, or ops leadInvestigates and changes systemsNothing else, if possible
CommunicationsWrites updates for customers and the companyScribe, on a small incident
ScribeKeeps the timeline: what was tried, and whenCommunications

The one combination to avoid is commander and fixer. A commander who is also debugging loses the wider view, and a debugger who is also coordinating does neither well. If the team is too small to split them, the commander should at least hand the technical work to the first available second pair of hands.

An example: the first ten minutes

This example is invented. At the start of a working day, the payments page starts returning errors. The on-call engineer gets paged, sees customers affected, and declares an incident within a couple of minutes without waiting to be sure of the cause. They open a channel and name a senior colleague as commander, then go back to investigating.

The commander asks three questions in the channel: who is looking at the last deploy, who is looking at the database, and who will write to customers. Three people answer with their names. The commander sets the rhythm: a short update every fifteen minutes, in the same place, even if it says “no change”. When the deploy turns out to be the cause, the commander asks the fixer whether a rollback is safe, hears yes, and says “roll back”. Ten minutes in, the business has one clear picture, the customers have had their first message, and nobody has been left to guess who is in charge.

When to declare

Teams wait too long to declare. The cost of a false alarm is a few minutes of attention. The cost of a late declaration is a tangle of people and a longer outage. A simple rule: declare an incident if any one of these is true.

  • Customers are, or are about to be, affected.
  • More than one person or team is needed.
  • The cause is not clear after a few minutes of looking.
  • Someone asks whether to declare.

Anyone on the team should be able to declare, without permission. Downgrading later is easy. Write the rule down, and tie it to the escalation policy so that a page and a declaration are connected.

Handover

An incident that runs long needs a change of commander. Tiredness is a risk to customers, because decisions get worse. Hand over deliberately:

  1. The outgoing commander states the current picture: what is wrong, the impact, what has been tried, what is being tried.
  2. They name the people in each role and any decision waiting.
  3. The incoming commander repeats it back, to be sure it is understood, and says so in the channel.
  4. The outgoing commander stays reachable for a short while, then leaves.

Say the handover out loud in the channel, so that everyone knows who is in charge. How to set up that channel is covered in running an incident channel in Slack or Teams.

A one-page role card

Print this, pin it, and give it to anyone who might take the role.

StepWhat to do
DeclareOpen the channel, state the severity and a one-line description
AssignName a fixer, a communications person and a scribe
Set a rhythmAgree when the next update comes, and keep to it
Keep the pictureMaintain a short summary: impact, status, next step
DecideMake the call on rollback, escalation and customer messages
Tell peopleMake sure customers and the company hear what they need to
Hand overPass the role on when tired or when the incident changes hands
CloseDeclare it resolved, record the follow-ups, schedule the review

After the incident, review how the coordination went alongside the technical cause. A blameless review, as described in AI-written postmortems, asks what made the role hard, and the answers improve the card. The numbers that show whether it works are in MTTA vs MTTR.

Frequently asked questions

What does an incident commander do?

They coordinate. The commander decides who works on what, keeps track of what is known and what is still a guess, makes the call when a decision is needed, and makes sure the right people outside the team are told. They do not fix the problem themselves, because the job needs attention on the whole incident.

Who should be the incident commander in a company of eight engineers?

A small group of experienced engineers who take the role in turn, with the on-call engineer acting as commander for the first minutes until a designated person takes over. The founder or the most senior person should not be the default: they are usually the one best placed to help fix the problem.

Can the incident commander also fix the problem?

On a small incident with one responder, one person may do both for a short time. On anything larger, split them, and hand the role over as soon as a second person is available.

When should an incident be declared?

Earlier than feels necessary. Declare when customers are affected, when more than one person or team is needed, when the cause is unclear after a few minutes, or when anyone is unsure. Declaring costs little and can be withdrawn. A late declaration costs time and confusion.

How do we train people for the role?

Let them shadow a commander, then take a small incident with the previous commander watching, and practise with a staged exercise. Give everyone the same one-page role card so the job is written down. Review each incident's coordination, not only its technical cause.