Skip to the protocols
Hush Hour Stress-reset protocols

Use whenThe first ten minutes of any outage, incident, failed deployment, or unexpected system change where more than one person is about to start working at once.

When the Systems Go Down

An outage produces a specific failure in the people responding to it: everyone starts doing something at once, the loudest problem gets the attention rather than the biggest one, and the first ten minutes produce three duplicate investigations. This protocol is eight minutes of ordering for the first person on the scene, and its main output is a single written line that says what is being worked on and by whom.

Work rhythm · 8 min · Published 5 October 2026 ·

Breathing rhythmOne rule only: three slow exhales before you type anything into a channel. Out-breath longer than in-breath. Everything else about this protocol is administrative.

The sequence

  1. 01

    Say what is known, out loud, once

    60 s

    One minute, one person: what is broken, since when, who is affected, what has been checked. Said aloud rather than typed, so that everyone hears the same version at the same time. Do not include theories in this minute.

  2. 02

    Write one line and pin it

    120 s

    Two minutes to put a single line where everyone can see it: what we are working on, who is on it, when the next update is due. Not a document, not a channel full of messages — one line, with a time in it.

  3. 03

    Split investigation from communication

    90 s

    Decide now who is not debugging: one person owns updates to customers, support, and management. Ninety seconds to name them, and they stop touching the system at that moment. The most common outage failure is the only person who understands it also being the only person answering questions.

  4. 04

    One hypothesis at a time, written down

    150 s

    Take the most likely cause, write it down, test it, write the result next to it, then move on. Two and a half minutes per hypothesis, and no parallel unlogged investigation — the log is what stops the same check being run three times by three people.

  5. 05

    Eat, drink, and set the next check-in

    60 s

    One minute, at the eight-minute mark: water, and a stated time for the next update even if it is twenty minutes away. A named time is what allows everyone else to stop refreshing the page, which is the difference between an outage that lasts two hours and one that eats the whole day.

The duplicate investigation problem

When four competent people start on an outage simultaneously, the usual result is that the same three checks are run three times and nobody runs the fourth. The fix is not coordination in general; it is a written log with a rule that no check is run without being written down first. Written hypotheses are also what make the post-incident review honest, because the review reads the sequence of what was believed at the time rather than what is obvious in hindsight.

Separating the fixer from the talker

In small teams the same person is often both, and it is exactly in small teams that this costs the most, because the fixer is interrupted every two minutes by a legitimate question from someone who cannot see the screen. Naming a communicator is ninety seconds of work and it is the single highest-leverage decision available in the first ten minutes. If there are only two of you, the communicator is the one who understands the system second-best, not the one with the least to do.

Stated update times

"We will update in twenty minutes" is worth more than "we will update soon", and it is not about accuracy — the second update can move the time. It is about giving twenty people permission to stop watching. An outage where nobody knows when the next information arrives is an outage where everybody refreshes everything, which produces a second, self-inflicted load problem on top of the first.

Environment tweaks

  • One place for the status line — a pinned message, a shared doc, a physical whiteboard. Two places means zero places.
  • Support gets told before customers; a support team that finds out from a customer will generate more noise than the outage itself.
  • Notifications for unrelated channels muted for the duration, not just deprioritised.

Questions people ask

What if I am the only person available?

Then the status line and the stated update time matter more, not less, and the communicator role is the one to give away first even if it is to someone who cannot fix anything. A colleague who can only say "we know, it is being worked on, next update at 14:20" removes a large share of the interruptions without needing any system access.

Should I start a call immediately?

A call helps when more than three people are involved or when the problem is not visible to everyone. Below that, a written status line is faster, because it does not require anyone to stop what they are doing and it leaves a record. If you do open a call, say at the start that it is for coordination and that debugging continues in parallel.

How does this connect to the post-incident review?

The written hypotheses and their results are most of the review's raw material. A review that has to reconstruct the timeline from memory produces a story about one heroic fix; a review that has the log produces a list of three checks that were run twice and one input nobody owned. The second kind changes something.

Informational only. This is a personal working habit, not an incident-management standard or security procedure. Follow your organisation's own escalation and notification requirements.