Defining SEV Levels for a Team of Five: When One Customer Down and Everyone Down Share the Same Pager

On a five-engineer rotation, every page lands on the same small pool of people. That is the constraint that makes severity definitions matter more for you than for a company with a dedicated incident-response org. A big-company SEV ladder assumes you can staff an incident commander, a communications lead, and an ops lead for every major incident. You cannot. So your SEV table has to do double duty: it decides how hard you respond, and it decides whether the on-call engineer wakes anyone else up.

The common failure mode is defining severity by who got paged rather than by measured impact. When one customer’s export job fails and the whole API is down, both arrive on the same phone. If your definitions do not distinguish them, the on-call engineer is left to improvise at 3 a.m. — which is exactly when improvisation is most expensive.

Severity is impact. Priority is urgency.

Atlassian draws a distinction that is worth borrowing even if you never use their tooling: severity is a measurement of impact, and priority is a measurement of urgency. A typo on the homepage is low severity and possibly high priority. An app crash affecting 0.05% of users is high severity and possibly not the top priority if something wider is also burning. The two fields answer different questions, and conflating them is how teams end up arguing about labels during an incident instead of resolving it.

For a small team, the practical consequence is that your SEV level should be a statement about measured impact — how many customers, which functionality, how much redundancy is gone — not a statement about how annoyed the on-call engineer is. Priority is the field you use to sequence work. Severity is the field you use to decide response intensity.

What a metric-driven ladder looks like

PagerDuty’s public incident-response documentation is a useful reference because it is specific. Their SEV-1 covers a critical issue warranting public notification and executive liaison, with functionality severely impaired for a large number of customers or a customer-data-exposing vulnerability. SEV-2 covers a critical system issue actively impacting many customers’ ability to use the product. Anything above SEV-3 is automatically considered a major incident and gets a more intensive response. SEV-3 covers partial loss of functionality not affecting the majority of customers, or something with the likelihood of becoming a SEV-2 if nothing is done. SEV-4 covers performance issues, individual host failure, delayed job failure, and cron failure that does not impact the event and notification pipeline. SEV-5 is cosmetic.

Two things in that ladder are worth copying directly. First, the definitions are metric-driven. PagerDuty explicitly recommends making your own definitions very specific, usually referring to a percentage of users or accounts affected. Second, the tie-breaker: if you are unsure whether an incident is SEV-2 or SEV-1, treat it as the higher one. During an incident is not the time to litigate severities; review the classification during the postmortem.

Atlassian’s ladder is similar in shape. Their SEV-1 examples include a customer-facing service down for all customers, a confidentiality or privacy breach, and customer data loss. SEV-2 includes a customer-facing service unavailable for a subset of customers and core functionality significantly impacted. SEV-3 includes a minor inconvenience with a workaround available and usable performance degradation. Atlassian also notes that SEV-3 incidents can be handled during working hours, while SEV-1 and SEV-2 generate an alert for on-call professionals regardless of time of day.

Adapting the ladder for five people

The mistake is to copy a five-tier ladder and assume it fits. It may not. Atlassian’s own guidance is that when setting severity levels you need to factor in the size of your tech team, your on-call schedules, your high- and low-traffic times, and the frequency of incidents. A five-person team with a single time zone and a service that is quiet between 2 a.m. and 7 a.m. has different constraints than a global service with follow-the-sun coverage.

A workable approach for a small team is to define three or four levels, each tied to a measurable condition, and to decide explicitly which levels page. A starting shape, adapted from the sources above:

  • SEV-1: Core functionality is unavailable for all or nearly all customers, or customer data is lost or exposed. Pages the on-call engineer and wakes a second person. Public or customer-facing communication is warranted.
  • SEV-2: Core functionality is significantly impaired for a subset of customers, or a critical internal pipeline is broken. Pages the on-call engineer. Escalation to a second person is at the on-call engineer’s discretion.
  • SEV-3: Partial loss of functionality not affecting the majority of customers, or a condition that will become SEV-2 if nothing is done. Pages during business hours; may page out of hours if the on-call engineer judges it necessary.
  • SEV-4: Performance degradation, single host failure, delayed job failure, or cron failure with no customer-visible impact. Ticket, not page.

The percentages and the specific functionality names are yours to fill in. The point is that each level has a written condition that a tired engineer can check against reality without a meeting.

The single-customer outage

This is the case the topic names, and it is where a metric-driven definition earns its keep. Suppose your service has 200 customers and one reports that a report export is failing. That is a partial loss of functionality for a small subset. Under a metric-driven definition it is a SEV-3 or SEV-4, not a SEV-1 — even though it arrived on the same pager as an everyone-down incident would.

The useful question is not “is this a SEV-1?” It is “does this meet our written threshold for a major incident?” If the answer is no, it can be handled as a SEV-3 or SEV-4 during business hours without waking the whole rotation. If the answer is yes — because the export is core functionality for a segment that represents a meaningful share of revenue, or because the failure is a symptom of something wider — then it escalates on the merits, not on the volume of the complaint.

This is also where the “assume the worst” tie-breaker helps a small team specifically. Debating severity during an incident has a social cost: someone has to argue that the thing they are being paged about is not actually that bad. That argument is easier to avoid than to win. Treating an uncertain incident as the higher severity and correcting it in the postmortem removes the debate from the moment when cognitive resources are scarcest.

Pair severity with postmortem triggers

Severity levels decide response intensity. They should also decide which incidents get a written postmortem. Google’s SRE book lists common postmortem triggers: user-visible downtime or degradation beyond a certain threshold, data loss of any kind, on-call engineer intervention such as a release rollback or traffic rerouting, a resolution time above some threshold, and a monitoring failure that implies manual incident discovery. The same chapter states that it is important to define postmortem criteria before an incident occurs so that everyone knows when a postmortem is necessary.

For a five-person team, the practical move is to write the postmortem trigger into the SEV table. If SEV-1 and SEV-2 always get a postmortem, and SEV-3 gets one when the on-call engineer intervened manually or when monitoring failed to catch it, then the team does not have to decide after the fact whether the incident “deserved” a writeup. The decision was made in advance.

This also gives you the correction mechanism for over-classification. If an incident was treated as SEV-2 under the assume-the-worst rule and the postmortem shows it was really SEV-3, that is a data point about your thresholds. It is not a failure. It is the system working as designed.

What the sources do not tell you

None of the primary sources cited here recommend a specific number of SEV levels for a five-person team, and none of them say a single-customer outage is always one level or another. That determination depends on your own metrics: how many customers or accounts you have, what counts as core functionality, and what your current redundancy posture is. The sources give you the shape of a metric-driven ladder and a tie-breaker rule. The numbers are yours.

What the sources do support is the underlying discipline. Google’s SRE book describes on-call engineers as expected to triage a page and work toward resolution, possibly involving other team members and escalating as needed, and notes that paging events take priority over almost every other task including project work. It also names clear escalation paths, well-defined incident-management procedures, and a blameless postmortem culture as the most important on-call resources. A written SEV table is one of those well-defined procedures. It is not bureaucracy; it is the thing that lets a small rotation respond consistently without a standing incident-command hierarchy.

The SRE workbook adds that formulating rules about how to communicate and coordinate before disaster strikes allows the team to concentrate on resolving an incident when it occurs, and lists declaring incidents early and often among the basic principles of incident response. A SEV table that is specific enough to apply in under a minute is what makes early declaration cheap.

A practical exercise

Before the next incident, write your SEV table with specific percentages and named functionality. Decide which levels page and which levels wait for business hours. Decide which levels trigger a postmortem. Then test the table against the last three incidents your team handled. For each one, ask: under these definitions, what level would it have been, and would the response have been different? If the answer is that the table would have produced a response the team could not actually sustain — waking three people on a Tuesday night for something that meets your SEV-2 threshold, for example — then the threshold is wrong, not the team.

Revise and repeat. The table is a living document, and the postmortem is where it gets corrected. For a related exercise in writing the recovery steps before you need them, see Write the Recovery Checklist Before You Need It.

FAQ

How many SEV levels should a five-person team have?

There is no sourced answer to this. Atlassian’s guidance is that you factor in team size, on-call schedules, traffic patterns, and incident frequency when setting levels. Three or four levels is a common starting shape, but the right number is the one your team can apply consistently under pressure.

Should a single-customer outage ever page?

It depends on your written threshold. If the affected functionality is core for a meaningful share of your customers, or if the failure suggests a wider problem, it may meet your major-incident threshold. If it is a partial loss for a small subset, a metric-driven definition likely places it at SEV-3 or SEV-4. The sources do not prescribe a specific mapping.

What if we disagree about the severity during an incident?

PagerDuty’s rule is to treat it as the higher severity and review during the postmortem. That removes the debate from the moment when it is most costly and puts the correction where it belongs.

Do we need a postmortem for every SEV-2?

Google’s SRE book lists common postmortem triggers including user-visible downtime beyond a threshold, data loss, on-call intervention, resolution time above a threshold, and monitoring failure. Writing your own triggers into the SEV table before an incident means the team does not have to decide afterward.

Is severity the same as priority?

No. Atlassian defines severity as a measurement of impact and priority as a measurement of urgency. They often align, but not always. Severity drives response intensity; priority drives sequencing.