Enterprise Operations · Incident Leadership · Judgement · 17 min read

CritSit Chaos: Where's the Runbook for That?

What happens when the documented process runs out, the evidence is incomplete and a room full of experts still needs to decide where to go next?

Joakim Domeij
By Joakim Domeij 3 October 2026 · 17 min read

Most organisations that operate critical systems have a process for major incidents. There are severity definitions, escalation paths, communication templates, incident bridges, support rotations and expectations for what happens afterwards. All of that is useful. Some of it is essential.

The problem is that the hardest parts of a critical incident rarely appear in the runbook.

Where does it tell you what to do when fifteen experienced people have spent two hours investigating and nobody knows where to look next? What does it say when two senior technical experts strongly disagree, the customer is becoming increasingly frustrated and someone still needs to decide where the investigation goes next? Where is the section explaining when to challenge an architect's assumption, when to let an engineer work in silence, when to wake another specialist at two in the morning, or when five minutes away from the bridge might actually get you closer to restoring production?

Those are the moments I have found most interesting about leading critical incidents. The process still matters, but eventually the runbook runs out. Judgement has to take over.

I have spent a lot of time on critical incident calls over the years, and the difficult ones rarely resemble the clean incident-management processes we document afterwards. Once a real production system is down, several things are happening at once. A customer is trying to understand the business impact while technical teams are trying to understand the failure. Executives want to know when service will be restored. Specialists are joining from different areas, sometimes in the middle of the night, with different information and different assumptions about where the problem might be. Someone needs to keep the investigation moving, keep the customer informed, make sure the right expertise is available and somehow prevent all that urgency from turning into noise.

I used to think the person leading that situation needed to be particularly good at finding answers. Experience has pushed me towards almost the opposite conclusion. You need enough technical understanding to follow what is happening, challenge assumptions and recognise where an investigation may need to go next, but you are unlikely to be the best database administrator, network engineer, developer, architect and business expert on the call. If you have ten or fifteen capable people assembled around a problem, trying to personally provide the technical answer would be a strange use of the expertise in the room. The harder job is creating the conditions in which those people can find the answer together.

Everyone joins the incident with assumptions, including me. If I am called because an application is failing, I already have an idea of the kind of problem I am walking into. A developer will naturally look at the application. A DBA will look at the database. Infrastructure will look at the platform. The customer will look at what their users can no longer do. None of that is wrong. Expertise gives people somewhere sensible to start. The danger begins when the starting assumption quietly becomes an established fact.

Perhaps everything initially points towards the database. The DBA investigates, but after twenty minutes there is nothing convincing there. That does not necessarily mean the database has been conclusively ruled out. It means it has moved down the list of likely explanations. The DBA can continue investigating in parallel while the shared attention on the bridge moves towards the integration layer and the relevant specialists are brought in. If someone asks me for an executive update at that point, I do not need to manufacture a breakthrough. I can explain that the database was our initial main suspect, that we have not found anything there that explains the issue, and that we are therefore shifting the main investigation towards another part of the system while the DBA continues working.

We have not fixed production, but we have still made progress because we know more than we did twenty minutes earlier. Troubleshooting is rarely a neat sequence where we investigate component A, permanently cross it off the list and move to component B. We are continually changing the relative likelihood of different explanations as new evidence appears. Something we deprioritised two hours ago may become interesting again because of something we discover later. The bridge's focus and the total investigation are not necessarily the same thing. Several useful investigations can continue in parallel, but having four technical conversations competing for attention on the same bridge rarely helps anyone.

There is also one very simple question that is surprisingly easy to forget when everyone is rushing to investigate the problem in front of them: Have we had this, or something similar, before? What did we do then?

The answer should not dictate the investigation. If somebody remembers an incident six months ago with similar symptoms that turned out to involve certificates, that does not mean we have another certificate problem. But it gives us something inexpensive to test. I do not need fifteen experts to stop what they are doing and investigate certificates. One person who knows where to look can check while the main investigation continues. Five minutes later they may have found something important, or they may simply have made that explanation less likely. Either result is useful.

That is one of the ways experience becomes valuable during incidents. The wrong lesson from experience is, “I have seen this before, so I know what it is.” The more useful lesson is, “I have seen something similar before, so there is another question worth asking.” Experience should give us better questions, not stronger assumptions.

The same principle applies to the information we receive at the beginning of an incident. A ticket may say the problem started at 14:10, but that may only be when somebody noticed it. A deployment may have happened at 13:42, errors may have started appearing at 13:47, a queue may have begun building at 13:55 and the first user may not have reported anything until 14:10. If we simply accept the ticket as the history of the incident, we may spend the next hour looking at the wrong period of time.

I like to establish the timeline early and keep refining it as we learn more. When did the system last behave normally? What changed after that? When did the first symptom appear? When did the customer notice it? What have we tried since then, and what actually happened after each action? I will often explain why I am asking: I am documenting this for the incident timeline and eventual RCA, so I want to distinguish what we have confirmed from what we currently believe.

That distinction between knowing and believing runs through almost every part of a critical incident. “Nothing changed today” can eventually turn into “there was no application deployment today, but there was a configuration change.” “All users are down” can turn into “all the users we have spoken to are affected.” “The network is fine” can turn out to mean that someone looked at a monitoring dashboard and saw nothing unusual. Each statement may have been made in good faith. The problem comes when it is repeated enough times that a qualified observation becomes an unquestioned fact.

The more pressure there is to appear certain, the more important it becomes to be precise about uncertainty.

This can be uncomfortable when senior people are involved. If an experienced architect says something is not the network, it is very easy for everyone else to move on. Sometimes that is absolutely the right call. Sometimes I still need to ask, “What have we seen that supports that?” or “Have we actually confirmed that, or is it still our working assumption?” That is not about challenging somebody's expertise for the sake of it. I need to understand what goes into the incident record, what I can tell the customer and how much confidence we should place in the next decision.

The same applies when someone asks for an estimated restoration time and we do not have enough evidence to provide one. An invented ETA may reduce pressure for ten minutes, but it creates a larger trust problem when the deadline passes and production is still unavailable. “I don't have enough evidence to give you a reliable restoration time yet” may not be the answer somebody wants, but sometimes it is the only accurate one.

There is an old idea that slow is fast and fast is slow, and few places demonstrate it better than a major production incident. The pressure to act can be enormous. Restart something. Roll it back. Change a configuration. Bring another team in. Give the customer an answer. Sometimes those actions are exactly right, but speed without structure can easily extend the incident. Restarting before capturing evidence can destroy useful information. Changing several variables simultaneously can make it impossible to know which change mattered. Pulling every available specialist onto the bridge can create more noise than expertise.

Taking five minutes to establish a timeline can save an hour. Taking thirty seconds to ask, “Do we know that, or are we assuming it?” can prevent a risky change. Slowing down briefly to understand the actual customer impact can completely change the way the incident should be prioritised.

That does not mean moving slowly. It means maintaining controlled urgency.

Understanding the impact is particularly important because the number of affected users does not necessarily tell us how serious an incident is. Ten reported users may mean ten users are affected, or simply that ten have reported it so far. One person being unable to use a system may be a minor issue, or that one person may be responsible for a process affecting thousands of others. A single device failing might normally be insignificant, but the consequences can look very different if that device happens to be broadcasting something to millions of people.

Impact is not simply how many people are affected. It is what happens because they are affected.

That means understanding what has actually stopped, what still works, whether there is a workaround, what financial or operational consequences exist, whether there is any data or regulatory exposure and whether the problem is spreading. Severity can legitimately change without the underlying technical fault changing at all because our understanding of its consequences has changed.

While the investigation continues, I also spend a surprising amount of time asking very simple questions. “What does that acronym mean?” “Can you explain what we're looking at?” “What does that actually mean for the customer?”

I have had technical people message me privately during incidents asking what something means because they do not want to ask what they think is a stupid question in front of a large group. Even when I know the answer, I may ask it aloud in a slightly different way. “Just so we're all using the same terminology, MTA means mid-term adjustment here, correct? A customer changing their address, for example?”

Ten seconds later the definition has been confirmed. Nobody has been embarrassed, and everyone on the call now has the same context. There may be fifteen people listening. Five understood the term immediately. Five had a reasonable idea. Five had no idea at all but did not want to interrupt the expert. Silence does not tell me which group they belong to.

Shared vocabulary is not the same thing as shared understanding.

This is also why I sometimes ask the person working on the strongest technical lead if it would help to share their screen and talk us through what they are seeing. It is not about fifteen people watching somebody type. Done at the right moment, it gives the room a shared view of the evidence. Someone else may notice something. A specialist from another area may connect what they see with something in their own domain. People learn from the explanation, and that learning changes the questions they ask next.

There is a balance. If talking through every action is slowing down the person doing the work, I would rather let them concentrate. The incident leader has to understand the rhythm of the incident: when to ask, when to summarise, when to challenge and when to stop talking.

That becomes even more important two or three hours into a difficult incident. People get tired. Someone who joined at midnight may still be there at three in the morning. People without a clear task naturally start switching off. Eventually you ask somebody a question because you think they are still on the call and get silence in return. Keeping people engaged does not mean constantly filling the air, but the room needs enough shared understanding that the people who are there can still contribute when the investigation changes direction.

Eventually, some incidents reach the point nobody wants to reach. We have investigated the areas that seemed most likely and we still do not have an explanation.

At that point, I think one of the worst things an incident leader can do is feel they personally need to invent the next answer. I would rather ask the room where we go from here. What haven't we looked at? What assumption might be wrong? Is there anything we investigated two hours ago that looks different now that we know more? Are we missing an answer, or are we missing the person who could see the answer?

Then give people space to think.

That last question can be particularly important. If everybody initially believed the problem was in the software, the people assembled on the call probably reflect that assumption. Perhaps nobody called the network engineer because there was no reason to wake them at two in the morning. After an hour of finding nothing in the software, the situation may look different. Our initial assumptions did not only determine where we looked. They determined who was in the room.

This is where some technical understanding helps the incident leader enormously. I do not need to know how to perform every investigation myself, but I need enough understanding to think one step ahead. If this test comes back clean, where are we likely to go next? Who would we need? Should I bring them onto the bridge now, or simply give them a heads-up so they are ready if we need them in fifteen minutes?

Being proactive does not always mean doing more. Sometimes it means quietly reducing the time required for the next decision.

The same judgement applies to restoration. At two in the morning, the technically perfect permanent fix is not always the right immediate objective. If we can safely restore service through a rollback, restart, failover or temporary workaround, that may be the responsible choice while the deeper investigation continues. On another incident, the workaround may introduce unacceptable risk to data integrity, security or stability and therefore be the wrong decision. There is no universal rule. Someone has to understand enough of the technical and business context to weigh the trade-off.

I also try to keep three very different promises separate: when I will communicate again, when I think we may be able to restore service and when we expect to have a permanent resolution. “I will update you again in twenty minutes” is not the same as “production will be restored in twenty minutes.” Mixing those expectations is an easy way to create unnecessary disappointment.

The communication itself changes depending on who needs it. Engineers may need to know which transactions are timing out and what the logs show. The customer may need to know which business processes are affected and what workaround exists. An executive may need to understand the scale of the impact, the current direction of the investigation, the risks and when the next meaningful update will arrive. The underlying truth should remain the same, but dumping every technical hypothesis onto every audience does not create transparency. It creates noise.

There is another system running alongside the production system during all of this: the group of humans trying to repair it.
That system can become unstable too.

People become frustrated. Customers may become angry, sometimes understandably so. Technical teams can disagree strongly. People can start defending their area rather than investigating the problem. A discussion about evidence can gradually become an argument about ownership. After several hours, capable individuals can become a considerably less capable team.

Sometimes the incident leader needs to calm the room down. That does not necessarily mean announcing that everybody needs to calm down, which is unlikely to improve anything. A natural pause may appear because a specialist needs ten minutes to download logs or run an investigation. That can be the perfect moment for a five-minute break. Other times, I may call a short break simply because continuing the conversation in its current state is becoming counterproductive.

Those five minutes can be useful in ways the wider bridge never sees. I can message one engineer privately and say that I think their point is valid, but we are getting stuck arguing about ownership and need to focus on what evidence would prove or disprove it. I can ask somebody else to confirm the timestamps while we are paused. I can speak to the customer contact and acknowledge the frustration while making it clear that I am going to reset the discussion around restoring service.

Then we come back and start again from what we know.

This is where trust matters enormously. If people have worked with me before, I have some credit. They know why I ask questions. They know I will communicate when I say I will. They know that if I ask them to focus on the investigation, I will deal with some of the surrounding pressure. They also know that I am not asking a simple question to expose somebody's lack of knowledge.

With a group I have never worked with, I do not automatically have that credit. My first few interventions matter. If every time an engineer becomes quiet for thirty seconds I ask for an update, I am not managing the incident. I am interrupting the person trying to solve it. If I ask questions that lead somewhere, keep communication predictable and protect useful technical work from unnecessary noise, trust starts to build.

A critical incident is therefore not simply a system that needs troubleshooting. For a few hours it creates a temporary team, often assembled from several departments, different companies, different levels of seniority and people who may never have worked together before. The incident leader has to monitor both systems: the production system that is failing and the human system trying to repair it.

The job is not to have every answer. It is to maintain the conditions under which good answers can still be found: enough structure to stop urgency becoming chaos, enough shared understanding for different experts to work on the same problem, enough curiosity to challenge assumptions, enough technical understanding to anticipate where the investigation may go next, and enough awareness of the room to know when to speak and when to get out of the way.

The Questions I Keep Coming Back To

These are not my attempt to write the missing runbook. A runbook cannot tell me exactly what the next decision should be when the evidence is incomplete, the people are tired and the situation is still changing. They are simply baseline questions I keep coming back to when the runbook no longer has the answer.

  • What is actually happening? What have we observed rather than simply been told?
  • Have we had this, or something similar, before? What happened then, and what did we do?
  • When did this actually start? Not when the ticket was raised, but the earliest evidence we can find.
  • What was the last known good state, and what changed afterwards?
  • What is the actual impact? Who or what is affected, what can they no longer do, and what happens because of that?
  • What still works? Where is the boundary between normal and failing behaviour?
  • What do we know, and what are we assuming?
  • What have we tried, and what actually happened after each action?
  • Where does the evidence point most strongly right now?
  • What have we deprioritised rather than genuinely eliminated?
  • What is cheap to check in parallel without distracting the main investigation?
  • What haven't we looked at?
  • Who are we missing? Is there expertise outside the room that could see the problem differently?
  • If our current hypothesis is wrong, where do we go next, and can we prepare for that now?
  • Can we safely restore service before we understand the permanent fix?
  • What does the customer need to know now, and when will they hear from us again?
  • What do executives need to understand to make any decisions required of them?
  • Does everyone understand the terminology and the problem we are currently discussing?
  • Is the room still helping the investigation, or are fatigue, frustration or disagreement starting to work against us?
  • Is our incident record separating what we observed, assumed, tested and confirmed?
  • Knowing everything we know now, does anything we looked at earlier need to be reconsidered?
  • If we still do not know the answer, what is the best question we can ask next?

I do not expect to have perfect answers to all of those questions. In some incidents, the most important answer is simply, “We don't know yet.” What matters is knowing what remains unanswered and creating enough structure around the uncertainty that the expertise in the room can continue working on it.

That is probably the part no runbook can completely capture. Leading a critical incident has never meant having all the answers for me. It means helping the people in the room find them without allowing pressure, assumptions, noise or frustration to make the problem harder than it already is.

Read Really, Another Leadership Book?

A practical examination of trust, judgement, ownership and leadership for the moments when the situation is more complicated than the advice.