A new service is about to take its first real traffic. For a team of two to fifteen engineers who are also the on-call rotation, the first incident on that service is usually a cold start: nobody has decided who gets paged, what they look at first, or which recovery path they attempt. A postmortem after that incident will produce useful learning, but the learning is paid for in user impact. A pre-mortem is the prospective mirror of that practice — same blameless ground rules, same insistence on owned and tracked action items, same one-page record, applied before the service is exposed rather than after it fails.
This is a proposed adaptation, not a documented industry practice. The postmortem craft it borrows from is well documented; the pre-mortem format below is a recommendation for lean teams that lack dedicated reliability coverage.
Why a pre-mortem, not another launch checklist
PagerDuty’s postmortem guide frames the after-the-fact practice plainly: performing postmortems after incidents is how teams learn what they are doing right, where they could improve, and how to avoid repeating mistakes. The same guide notes that a successful postmortem process rests on a culture of honesty, learning, and accountability, and that culture change can be led from any role even though it benefits from management buy-in.
A pre-mortem applies that same posture to a service that has not yet failed. The Google SRE book states that for a postmortem to be truly blameless it must focus on identifying contributing causes without indicting any individual or team, and that a blamelessly written postmortem assumes everyone involved had good intentions and did the right thing with the information they had. Those two sentences are the entire ground rule for the pre-mortem. The team is not predicting who will break the service; it is enumerating the conditions under which the service, as designed, will fail, and deciding in advance what the response will be.
The Google SRE book also notes that it is important to define postmortem criteria before an incident occurs so that everyone knows when a postmortem is necessary. The pre-mortem extends that logic one step earlier: define the failure modes and the response criteria before the service takes traffic, so the first incident is a rehearsal of a known plan rather than an improvised one.
The 45-minute bound is a design constraint
Forty-five minutes is not a target to fill. It is a constraint that forces the team to cluster and rank failure modes instead of enumerating every conceivable one. A session that runs ninety minutes will be scheduled once and then skipped. A session that runs forty-five minutes can be run before every new service, not only before the largest launches.
The agenda below assumes four to eight participants. One person facilitates and one person writes. The facilitator does not need to be the most senior engineer; the writer does not need to be the service owner.
Minutes 0–5: framing and scope
State the service, the traffic it is about to receive, the data it will touch, and the people who will be paged when it fails. Name the blast radius explicitly: which other services depend on this one, and which users notice. Write the scope on the shared document so it does not drift.
Minutes 5–15: silent failure-mode generation
Everyone writes failure modes independently, in silence, into the shared document. Silent generation prevents the first confident voice from anchoring the room. The prompts below are the ones that fit lean teams; use them as starting points, not as a fixed list.
- What breaks at 03:00 when only one person is awake?
- What happens if the person who built the service is on holiday?
- What does the first bad deploy look like?
- What does the first bad migration look like?
- What does the first credential rotation look like?
- What does the first cold cache look like under real traffic?
- What does the first connection-pool exhaustion look like?
- What does the first missing index look like on the largest table?
- What does the first rate-limit or auth surprise look like at the edge?
Minutes 15–30: clustering and ranking
Group the failure modes into clusters. Rank each cluster by likelihood and blast radius. The ranking does not need to be precise; it needs to be explicit enough that the team agrees on which three or four clusters deserve action items this week and which can be deferred. A cluster that survives ranking should produce at least one action item.
Minutes 30–40: owners and tracking
Each action item gets a single owner, a priority, and a tracking artifact. The Google SRE workbook states that action items without clear owners are less likely to be resolved, and that it is better to have a single owner and multiple collaborators. The same workbook states that without a formal tracking process, action items from postmortems are often forgotten, resulting in outages. The pre-mortem inherits both rules. Vague verbs like “improve” or “make better” are not action items; they are intentions. An action item names a specific change, a specific owner, and a specific place where its completion is visible.
The workbook also states that trying to change human behavior is generally less reliable than changing automated systems and processes. When a failure mode resolves to “the on-call engineer should remember to do X,” the action item should be rewritten to make X automatic, or to make the omission visible in monitoring, rather than to ask a tired human to remember.
Minutes 40–45: the one-page record
Write the record while the session is still open. The record contains: scope, participants, failure modes considered, action items with owners and tracking, and a date to revisit. The Google SRE workbook states that a postmortem is a factual artifact that should be free from personal judgments and subjective language, should consider multiple perspectives, and should be respectful of others. The pre-mortem record follows the same standard. It is a document, not a transcript of who said what.
Turning imagined failures into rehearsals
An action item that only changes a document is cheap and often sufficient. An action item that changes what the team has actually done before launch is stronger. For each failure mode that survives ranking, ask two questions: which signal would reveal it, and which rehearsal would exercise the recovery path?
Consider a hypothetical team about to expose a new service to real traffic. One failure mode they might imagine is that the only engineer who knows how to restore the primary database is unreachable at 03:00. The action item is not a note in a wiki; it is a restore lottery scheduled before launch, in which a different engineer performs the restore from a recent snapshot on a scratch host. The restic documentation describes the restore command and the use of the word latest to select the most recent snapshot, and it warns that restoring data in-place can leave files in a partially restored state if the operation is interrupted, recommending a current backup before restoring a different snapshot. A restore lottery that runs against a scratch host, not against production, is the rehearsal that turns that warning into a known procedure.
Another failure mode a team might imagine is that the first bad migration leaves the database in a state that requires point-in-time recovery. The PostgreSQL documentation states that PostgreSQL maintains a write ahead log in the pg_wal/ subdirectory of the cluster’s data directory, and that the log records every change made to the database’s data files. It states that continuous archiving supports point-in-time recovery, making it possible to restore the database to its state at any time since the base backup was taken. It also states that to recover successfully using continuous archiving, you need a continuous sequence of archived WAL files that extends back at least as far as the start time of your backup, and that you should set up and test your procedure for archiving WAL files before you take your first base backup. The pre-mortem action item is therefore concrete: verify before launch that WAL archiving is running, that a base backup exists, and that at least one engineer has performed a PITR restore on a non-production copy.
The same documentation notes that if archiving falls significantly behind, the amount of data that would be lost in a disaster increases, and the pg_wal/ directory will contain large numbers of not-yet-archived segment files that could eventually exceed available disk space. It also notes that the archive command is only invoked on completed WAL segments, and that archive_timeout can force the server to switch to a new segment at least that often. A pre-mortem that names “WAL archiving silently falls behind” as a failure mode should produce an alert on archive lag and a documented archive_timeout value, not a reminder to check the directory.
A third hypothetical failure mode: the first credential rotation locks out the deploy pipeline. The action item is a one-page runbook that names the break-glass path and the person who holds the second key. The runbook is written before launch, not during the incident. The Recovery Checklist Before You Need It is a useful companion here, because it forces the team to write the recovery steps down while the service is still healthy.
What the pre-mortem is not
The pre-mortem is not a substitute for the postmortem. The Google SRE workbook states that when written well, acted upon, and widely shared, postmortems can be a very effective tool for driving positive organizational change and preventing repeat outages. The pre-mortem reduces the number of surprises; the postmortem after the first real incident remains the place where the team learns what the rehearsal missed. The two practices share a repository, a tone, and a standard for action items, but they answer different questions.
The pre-mortem is also not a launch coordination checklist, a canary release, or a rollback plan. It does not replace any of those. It sits alongside them and produces the action items that make them specific to this service.
Nor is the pre-mortem a prediction. The team is not trying to name the exact failure that will occur. It is pre-deciding who will be paged, what they will look at first, and which recovery path they will attempt. The value is in the pre-decision, not in the accuracy of the forecast.
Closing the loop
Schedule the review before the session ends. The Google SRE workbook states that the value of a postmortem is proportional to the learning it creates, and that the more people who can learn from past incidents, the less likely they are to be repeated. The pre-mortem record should be shared as widely as the team’s postmortems are shared, and it should live in the same repository. When the first real incident occurs, the team compares what happened against what was anticipated. The gap between the two is the input to the next pre-mortem.
For a team that is its own on-call rotation, the pre-mortem is the cheapest rehearsal available. It costs forty-five minutes and a one-page document. It does not require a dedicated SRE function, a new tool, or a vendor. It requires only that the team treat the first real traffic as a rehearsal of a known plan rather than a cold start.
Questions readers ask
How is a pre-mortem different from a design review?
A design review evaluates whether the service is built correctly. A pre-mortem assumes the service is built as designed and asks how it will fail under real traffic, who will be paged, and what the response will be. The two meetings can be adjacent, but they produce different artifacts.
Do we need a facilitator who is not on the team?
No. The facilitator’s job is to keep the timebox and to prevent the first confident voice from anchoring the room. Any team member can do it. The writer’s job is to keep the record factual and free of personal judgment, which is the same standard the Google SRE workbook applies to postmortems.
What if we cannot rank failure modes by likelihood?
Rank by blast radius instead. The point of ranking is to decide which clusters get action items this week. A cluster that would take down the service for every user outranks a cluster that would degrade one internal dashboard, even if the likelihood is unknown.
How many action items should come out of a 45-minute session?
Enough to cover the clusters that survived ranking, and no more. The Google SRE workbook states that action items without clear owners are less likely to be resolved. A short list with single owners and tracking artifacts is more useful than a long list of intentions.
Should the pre-mortem record be public?
Share it as widely as the team’s postmortems are shared. The Google SRE workbook states that the value of a postmortem is proportional to the learning it creates. The same logic applies to the pre-mortem record, with the caveat that it may contain details about unreleased services that should not leave the organization.