- Joel Zamboni

Incident Management for SaaS: Runbooks, On-Call and Post-Mortems

A practical incident management guide for SaaS teams: severity levels, an on-call rotation a small team can sustain, runbooks, updates and blameless reviews.

Every SaaS product will have incidents. The difference between a bad night and a lost customer is rarely the bug itself; it is how quickly someone noticed, whether they knew what to do, and how clearly you communicated while it was happening. Incident management is the set of habits that makes those things predictable: a shared definition of severity, an on-call rotation that does not burn people out, runbooks that answer the first questions, a plan for communication, and post-mortems that make the next incident less likely. This guide covers each one for a small team running software in production, not security operations or emergency services.

What incident management covers: the five steps

An incident is any unplanned event that degrades your service for users or puts it at risk. Software incident management is the process from the moment something goes wrong to the moment you have learned from it:

  1. Detect: monitoring or a customer tells you something is wrong.
  2. Triage: decide how bad it is and who needs to be involved.
  3. Respond: mitigate first, then fix.
  4. Communicate: keep your team, leadership and customers informed.
  5. Learn: write the post-mortem and follow through on the actions.

Google’s SRE book suggests a useful rule for when to formally declare an incident: if you need a second team, if customers can see the problem, or if it is still unresolved after an hour of focused work. Declaring early costs little. Declaring late means the response is improvised.

Define severity levels before you need them

Severity levels decide who gets woken up and how fast. Agree on them while things are calm, write them down, and keep them few. A starting point for a small SaaS team, with incident management examples for each level:

LevelMeaningExampleResponse
SEV-1Product down or data at risk for many customersLogin fails for everyone; data lossPage immediately, all hands, customer updates
SEV-2Major feature broken or badly degradedCheckout errors for a large share of usersPage on-call, updates to affected customers
SEV-3Partial degradation with a workaroundOne integration delayed; slow reportsNext business hours, ticket
SEV-4Minor issue, no real user impactCosmetic bug; noisy alertBacklog

PagerDuty publishes its own severity level definitions, which are a good reference if you want more detail. It also offers the most useful single rule in this area: if you are unsure between two levels, treat it as the higher one and settle the classification in the post-mortem, not during the incident.

Designing an on-call rotation for a small team

The on-call rotation is where incident response plans meet real people, and where small teams most often get it wrong. The usual failure is that the CTO or one senior engineer is always on call because nobody else knows the system.

Google’s guidance gives a sense of scale. Its SRE workbook says sustainable 24/7 coverage needs at least eight engineers at a single site, or five per site across two sites, and targets no more than two incidents per 12-hour shift. Most startups have nowhere near that headcount for operations, so they have to design around it.

Principles that work for small teams:

  • Use primary and secondary. The primary responds; the secondary is the backup if the primary does not acknowledge within a set time. Escalation should be automatic, not a phone tree.
  • Rotate weekly and hand off on purpose. A short handoff note covering open issues, recent deploys and anything flaky saves the next person an hour.
  • Page only for things that need a human now. Every page should be actionable and tied to user impact. Everything else goes to a channel or ticket for business hours. Alert fatigue is how real pages get ignored.
  • Protect the person on call. Lighter project work during the on-call week, time off after a bad night, and compensation for out-of-hours work. Google’s workbook says out-of-hours work should be compensated.
  • Track the load. Count pages per week and per shift. If one service generates most of them, fixing it is the best reliability work you can do.

If you cannot staff a fair rotation internally, that is a signal to share the load with an outside team rather than spread it thinner across your developers.

Runbooks: answer the first ten minutes

A runbook is a short document attached to an alert that tells the responder what the alert means and what to check first. Good runbooks turn a 3am page from a research project into a checklist.

Every alert should link to a runbook with:

  • What the alert means in plain language and why it matters to users.
  • How to confirm it: the dashboard, the query, the log search.
  • Likely causes, starting with the most common: a recent deploy, a dependency, capacity.
  • Safe first actions: roll back, scale out, fail over, restart, with the exact commands.
  • When to escalate and to whom.

Keep runbooks in version control next to the code or infrastructure they describe, and review them after every incident that used one. An alert with no runbook is unfinished work. A runbook that has not been touched in a year is probably wrong.

Incident management roles and responsibilities

Responders fix the system; someone else should handle the talking. Google’s incident model separates the response into four roles: incident command, which holds the overall state and directs the response; operations, which applies the fixes; communication, which keeps stakeholders updated; and planning, which handles longer-term work such as filing bugs and arranging handoffs. In a small team one person may hold two of them, but naming who is in charge and who is updating stakeholders avoids the classic failure where everyone debugs and nobody tells customers anything.

Communication during an incident

Practical habits:

  • One channel per incident, so the timeline is in one place.
  • Regular updates on a fixed schedule, even when the update is “still investigating.” For a SEV-1, every 30 minutes is a reasonable default.
  • A public status page for customer-facing incidents, with templates written in advance.
  • Plain language. Say what users are experiencing, what you are doing and when the next update will come. Avoid guessing at causes in public.

Blameless post-mortems

After the incident, write it up. Google defines a blameless post-mortem as one that focuses “on identifying the contributing causes of the incident without indicting any individual or team for bad or inappropriate behavior.” The assumption is that people made reasonable decisions with the information they had. If someone could take production down with one command, the problem is the system that allowed it.

Decide in advance which incidents get a post-mortem. Google’s list is a good start: user-visible downtime beyond a threshold, any data loss, an on-call intervention such as a rollback, a long time to resolve, or a monitoring failure where a person found the problem first.

A useful post-mortem has:

  1. Summary and impact: what happened, who was affected, for how long.
  2. Timeline: detection, key decisions, mitigation, resolution.
  3. Contributing causes: usually several, technical and process.
  4. What went well and what was lucky.
  5. Action items with an owner and a date.

The action items are the only part that changes anything. Track them like any other engineering work and review the open ones monthly. A post-mortem whose actions never ship is a diary entry.

A minimal incident response plan

If you have nothing today, this is enough to start:

  • Four severity levels, written down.
  • A primary and secondary on-call with automatic escalation.
  • Alerts on user-facing symptoms, each linked to a runbook.
  • An incident channel template and a status page.
  • A post-mortem template and a rule for when to use it.

Each of those can be set up in a week. Improve them after every real incident.

Frequently asked questions

What do you mean by incident management?

Incident management is the process a team follows from the moment a service degrades or is at risk until it has learned from the event. For a SaaS team it covers severity levels, on-call, runbooks, communication with customers and post-mortems. The goal is to make the response predictable instead of improvised.

What are the 5 steps of incident management?

A practical version for software teams is detect, triage, respond, communicate and learn. You notice the problem through monitoring or a customer, decide how severe it is, mitigate before fixing, keep people informed throughout, and finish with a blameless post-mortem whose action items actually ship.

Who does what during an incident?

Google’s model names four roles: incident command, operations, communication and planning. The incident commander coordinates and makes the calls, operations works on the system, communication handles updates, and planning takes care of follow-up work. On a small team, one person often covers two roles, as long as everyone knows who is in charge.

Where agents and engineers fit

Much of incident management is watching, routing and gathering context, which is work machines do well and humans do badly at 3am. Judgment, fixes and customer communication are human work.

At Webera, Sentinel builds the monitoring, detects anomalies and creates alert rules with runbooks. Dispatcher designs on-call schedules and escalation policies, routes incidents by severity and service ownership, filters noisy alerts and creates incident response playbooks. Warden checks that every alert has a runbook before 24/7 monitoring goes live. Then a senior engineer on call reads the context, makes the call and ships the fix. AI is good at the first analysis and still not reliable at finding root causes alone, a point we cover in site reliability engineering vs. DevOps. Incidents and cost often meet, too: aggressive rightsizing is a common cause of both, as our AWS cost optimization playbook explains.

Take the pager off your developers

Webera is a DevOps subscription for SaaS teams on AWS. Every plan includes 24/7 monitoring and incident response, unlimited requests and all nine agents. Pro is $4,999 a month, Elite $9,999 with a named engineer, and Platform $14,999 with a team inside your sprints. Each has a three-month minimum, then month to month, and you own all the code and configuration.

If your CTO is still the on-call rotation, compare the plans or read what a managed DevOps subscription replaces.

Need DevOps expertise?

Our team of senior engineers can help you implement these practices.

Book a Discovery Call