How to Run a Fire Drill: Incident Response for SMBs

Running a tabletop fire drill gives your team a practiced, role-based response before the first real incident. Roles are defined, the runbook is tested, and no one is figuring things out under pressure when something breaks. Setting one up takes a couple of hours and a one-page document.

This post covers:

Verified against PagerDuty Incident Response documentation (response.pagerduty.com) and Atlassian Incident Management guides (atlassian.com/incident-management) on 2026-06-29.

Prerequisites

Step 1: Define What Counts as an Incident

Not every support request or technical hiccup is an incident. Defining a threshold before the drill means the team knows what triggers the response process and what goes through normal channels.

A working definition for most SMBs: an incident is any unplanned disruption that is affecting customers, blocking revenue-generating work, or creating meaningful data or security risk.

From there, assign two severity levels:

Write these down before the tabletop. Every subsequent step references them, and the team needs a shared definition of what P1 versus P2 looks like in your specific environment.

Step 2: Assign Three Roles

Incident response works with clear role separation. Three roles cover most scenarios:

Incident Lead. Owns the incident from detection to resolution. Coordinates the response, makes decisions, and drives the postmortem. The Incident Lead is not necessarily the person fixing the issue.

Communicator. Handles all external and internal communication during the incident: status updates to stakeholders, customer-facing messaging when needed, and progress updates to the team. Keeps the Incident Lead focused on resolution, not messaging.

Resolver. The technical owner who diagnoses and fixes the issue. Often an engineer or whoever manages the affected system.

At a smaller team, one person may hold two roles in a real incident. That is fine. What matters is that someone has each responsibility and knows it before the incident starts. The drill is where that gets tested.

Step 3: Build the One-Page Runbook

The runbook is the response template the team follows during an incident. One page. It answers who does what, in what order, and who gets notified.

Structure it as:

  1. Detection. How incidents get reported. (Example: alert in a dedicated Slack channel, customer support escalation, or a monitoring tool notification.)
  2. Assessment. Who determines severity. (Example: Incident Lead assesses within 15 minutes of the initial report.)
  3. Activation. How roles are assigned. (Example: Incident Lead notifies the Resolver and Communicator via direct message.)
  4. Communication cadence. How often updates go out. (Example: internal update every 30 minutes; external customer update at 60-minute intervals for P1s.)
  5. Resolution. How the Incident Lead declares the incident closed.
  6. Postmortem. When and how the team reviews what happened. (Example: within five business days for P1s, within ten for P2s.)

Add a contact block at the bottom: names, Slack handles, and phone numbers for each default role assignment. The runbook should be usable when people are not in front of their computers.

Step 4: Run the Tabletop Exercise

A tabletop exercise is a facilitated walkthrough of a simulated incident scenario. No live systems are touched. The team talks through what they would do, step by step.

Pick a scenario that is realistic for your stack. Useful starting points: a key SaaS tool goes down during business hours; a customer reports they cannot access their account; a third-party integration fails mid-process.

Before the session:

During the session (60 minutes):

  1. Facilitator presents the scenario at t=0
  2. Team works through the runbook step by step, narrating what they would actually do: who they contact, what message they send, what they check first
  3. Facilitator introduces two or three injects midway: complications that change the picture (examples: the Resolver is unavailable; the outage is now affecting a second system; a customer is escalating publicly)
  4. Team continues narrating through the runbook with each new inject

The goal is not to resolve the scenario. It is to surface the gaps in the process. Log every moment where the team pauses, disagrees, or has to improvise. Those pauses are where the runbook breaks down.

Step 5: Debrief and Update the Runbook

Within 24 hours of the tabletop, run a 30-minute debrief:

Update the runbook based on the gaps the drill surfaced. If a role was unclear, sharpen the definition. If a step assumed something that would not be available in a real incident, fix the assumption. If the exercise revealed a contact or tool that should be in the runbook but is not, add it.

Add a version date to the runbook: "Last reviewed: 2026-06-29." That date tells the team when it was last verified against actual conditions, and when it is time for the next review.

Step 6: Schedule the Next Drill

A runbook that is not rehearsed drifts out of date as the team and stack change. Set the next drill date before this session ends.

Quarterly is the right cadence for most SMBs with active development or customer-facing systems. Twice a year works for teams with stable stacks and low turnover.

Rotate the scenario each time. Running the same type of incident every drill trains the team for one event type. Cycle through: SaaS outage, security incident, integration failure, and data issue. Each one tests a different section of the runbook.

Verify

The practice is working if:

If any of these are missing, run a shorter 30-minute refresher rather than waiting for the next scheduled session.

Troubleshooting

The team treats the tabletop as a pass/fail exercise. Reframe it before the session starts. The point is to find what breaks. A drill that reveals no gaps was not realistic enough.

The Incident Lead keeps defaulting to the Resolver role during the exercise. This is the most common failure mode. The Incident Lead's job is to coordinate, not fix. In the next drill, assign the Incident Lead role to someone who does not have technical access to resolve the scenario. This forces the coordination role to stay separate.

The runbook is not being followed in real incidents. The most common cause is that the runbook lives somewhere the team does not find under pressure. Pin the runbook link in the incident response Slack channel and the team's internal wiki. It should be findable in under 30 seconds.

Roles are not being assigned when a real incident starts. Add an explicit activation step at the top of the runbook: the first person to identify a potential P1 or P2 posts in the incident channel and tags the Incident Lead. That tag is the activation trigger.

Ready to Build Your Incident Response Process?

ScaleIt helps startups and SMBs build operational processes, including incident response runbooks, team training, and escalation frameworks, as part of our Operations Blueprint and Startup Toolkit Launch engagements. Book a free call to talk through what your team needs.

Verified against PagerDuty Incident Response documentation (response.pagerduty.com) and Atlassian Incident Management guides (atlassian.com/incident-management) on 2026-06-29.