Most teams can state their backup schedule. Fewer can state their recovery point objective as a measured interval. The schedule is a claim; the RPO is the number you get when you actually restore and count how many minutes of committed work disappeared. For a lean team running its own on-call rotation, the gap between those two numbers is where incidents get worse than they needed to be.

This is a procedure for closing that gap. It uses PostgreSQL continuous archiving as the worked example because the mechanics are documented precisely, but the method applies to any system where you can name the last recoverable transaction and compare it to the moment of failure.

RPO is an interval, not a schedule

A nightly base backup at 02:00 tells you when the backup started. It does not tell you how much committed data you lose if the primary dies at 14:37. That number depends on two things: how far the archived write-ahead log (WAL) chain reaches, and how long the restore itself takes. The PostgreSQL documentation is direct about the first: point-in-time recovery lets you restore the database “to its state at any time since your base backup was taken.” The second is wall-clock time you have to measure, not estimate.

So the real RPO is the larger of two intervals:

  • Archive lag — the gap between the last committed transaction and the last WAL segment safely stored in archive storage.
  • Restore time — the wall-clock duration from “we decide to restore” to “the database is serving reads and writes again.”

If archive lag is 40 minutes and restore takes 25 minutes, your real RPO is not 40 minutes. It is 65 minutes, because the data you can recover is already 40 minutes stale before you start, and you spend another 25 minutes getting it back online.

Why the archive lag is usually worse than you think

PostgreSQL invokes the archive command only on completed WAL segments. A segment is normally 16MB. On a busy system, segments fill quickly and archive lag stays small. On a quiet system — nights, weekends, low-traffic services — a segment can sit partially full for a long time. The documentation states the consequence plainly: “if your server generates only little WAL traffic (or has slack periods where it does so), there could be a long delay between the completion of a transaction and its safe recording in archive storage.”

That delay is your archive lag. It is invisible in normal operation because nothing is failing. It becomes visible the moment you need to restore.

The documented lever is archive_timeout. Setting it forces the server to switch to a new WAL segment at least that often, which bounds how old unarchived data can be. The docs note that a setting of “a minute or so” is usually reasonable, and warn that very short values bloat archive storage because forced-switch segments are still full-length files. You can also force a switch manually with pg_switch_wal when you want a just-finished transaction archived immediately.

There is a second constraint that matters for the lottery: to recover using continuous archiving, you need a continuous sequence of archived WAL files extending back at least as far as the start time of your base backup. If the chain is broken — a missing segment, a failed archive command that was never retried, a retention policy that deleted a segment you still needed — the restore stops at the break. The lottery will find this.

The restore lottery: pick a minute, restore to it, measure the gap

The drill is simple to describe and deliberately inconvenient to run. That is the point.

  1. Pick a random past minute. Not a convenient one. Not the end of a quiet window. Use a timestamp from the middle of a normal working period, ideally one you did not choose in advance. Write it down before you start.
  2. Restore the base backup to a scratch environment. This is the file-system-level backup plus the WAL chain. pgBackRest’s user guide defines a restore as requiring “the backup files and one or more WAL segments in order to work correctly,” which is the same shape whether you use pgBackRest, pg_basebackup plus manual WAL shipping, or another tool.
  3. Replay WAL up to your chosen timestamp. In PostgreSQL this is the recovery_target_time setting in the recovery configuration. The server replays archived segments until it reaches the target, then stops.
  4. Record the timestamp of the last recoverable transaction. This is the number that matters. Compare it to your chosen target. The difference is your measured archive lag for that moment.
  5. Record the wall-clock restore time. From “start restore” to “database accepts connections.” This is the second half of your real RPO.
  6. Note which segment boundary caused the gap. If the last recoverable transaction is 38 minutes before your target, and the next archived segment starts 38 minutes later, you have found the boundary. That is a configuration fact, not a mystery.

Do this on a schedule — monthly is a reasonable starting cadence for a small team — and vary the target time. A drill that always restores to the same convenient minute proves the backup works. It does not measure the RPO.

A worked example

Suppose a team targets a 5-minute RPO. They run a lottery against a quiet Sunday window and pick 03:17 as the target. The restore completes. The last recoverable transaction is at 02:39. The gap is 38 minutes.

Nothing failed. The backup restored correctly. The WAL chain was intact. But the last archived segment before 03:17 was completed at 02:39, and no further segment filled until 03:22 because the system was idle. The archive command had nothing to archive. The measured RPO for that window is 38 minutes, not 5.

The fix is not a new backup tool. It is archive_timeout = 60s, which forces a segment switch every minute and bounds the lag. The team re-runs the lottery the following week against a different quiet window and measures again. This time the gap is under two minutes. The restore time is 22 minutes. The real RPO is now roughly 24 minutes — still worse than the 5-minute target, but now the team knows why and can decide whether to invest in faster restore or accept the number.

That example is illustrative, not a reported incident. The point is the shape of the finding: the drill, not the backup tool, revealed the gap.

What the lottery measures that a backup test does not

A backup test answers “can I restore?” A restore lottery answers “how much do I lose, and how long does it take?” Those are different questions with different failure modes.

The lottery will surface:

  • Archive lag during quiet periods. The completed-segment constraint means low-traffic windows accumulate lag. archive_timeout is the documented bound.
  • Broken WAL chains. If a segment is missing, the restore stops early. The lottery finds the break because it tries to reach a specific timestamp.
  • Retention that deletes too aggressively. If your retention policy removes WAL segments or base backups before the next one is verified, the recoverable window shrinks. pgBackRest’s documentation notes that incremental backups depend on prior backups being valid to restore, so a retention policy that removes a full backup can invalidate a chain of incrementals.
  • Restore time that nobody measured. Teams often know the backup size but not the restore duration. The lottery produces a wall-clock number.
  • Object storage versioning gaps. If backup artifacts live in object storage, versioning behavior affects what you can recover. Google Cloud Storage’s Object Versioning documentation explains that a noncurrent version is retained each time a live object is replaced or deleted, and recommends soft delete over Object Versioning for protection against permanent data loss from accidental or malicious deletion. The lottery should confirm that the version you expect to restore is actually the version you get.

Turning the number into a decision

The measured RPO is only useful if it changes something. Three outcomes are common:

The measured RPO is worse than the target, and the gap is archive lag. Shorten archive_timeout. The PostgreSQL docs note that a minute or so is usually reasonable and that very short values bloat archive storage. Test the tradeoff. Re-run the lottery.

The measured RPO is worse than the target, and the gap is restore time. The fix is usually operational: faster storage, a pre-staged base backup, a documented restore procedure that the on-call engineer can follow without re-deriving it. This is where a written recovery checklist earns its keep — see Write the Recovery Checklist Before You Need It for the shape of that document.

The measured RPO is worse than the target, and the team decides to accept it. That is a legitimate outcome, but it should be written down. “We accept up to 40 minutes of data loss during quiet windows because the cost of eliminating it exceeds the value” is a decision. Leaving it unstated is not.

In all three cases, the measured number and the drill procedure go into the runbook. The next on-call engineer should be able to reproduce the measurement without re-deriving the method. That is how the number stays honest as the system changes.

What the lottery does not tell you

The lottery measures recoverability of the database. It does not measure:

  • Whether your application can tolerate the restored state. A database restored to 03:17 may be consistent, but the application may have written external state — queue messages, object storage keys, third-party API calls — that the database does not know about.
  • Whether your encryption keys are available. If backup artifacts are encrypted, the restore depends on key access. The lottery should include a key-retrieval step, or it is measuring a restore that cannot happen.
  • Whether your monitoring would have told you the archive was lagging. The lottery is a drill; monitoring is the production control. If archive lag can grow to 38 minutes without an alert, the lottery will keep finding the same gap.

These are separate measurements. The lottery is the one that produces the RPO number.

Frequently asked questions

How often should we run a restore lottery?

Monthly is a reasonable starting cadence for a team of 2–15 engineers. The goal is to catch drift: configuration changes, retention policy changes, schema growth that slows restore, and quiet periods that expose archive lag. If your system changes frequently, run it more often. If it is stable, quarterly may be enough — but the first few runs should be close together so you can see whether the number is stable.

Can we run the lottery against production?

No. Restore to a scratch environment. The lottery is a measurement, not a failover. Running it against production converts a drill into an incident.

What if we use restic instead of pgBackRest?

The lottery method applies to any backup system where you can name the last recoverable point. For restic, the snapshot timestamp is the starting point. The restic documentation covers snapshot listing, restore, and repository integrity checking. The measurement is the same: pick a target time, restore, record the timestamp of the last recoverable data, and compare. The difference is that restic snapshots are point-in-time by nature, so the “archive lag” is the interval between snapshots rather than the interval between WAL segments.

What if the restore fails?

That is a successful lottery. A failed restore in a drill is a failed restore you did not have during an incident. Record what failed, fix it, and re-run. The PostgreSQL documentation notes that archive commands should refuse to overwrite pre-existing archive files, and that a nonzero exit status tells PostgreSQL the file was not archived and it will retry periodically. If your archive command has been silently failing, the lottery will find the gap in the WAL chain.

Do we need a dedicated SRE to run this?

No. The procedure is a sequence of documented commands and a stopwatch. The value is in running it on a schedule and writing down the number. A team that already runs its own on-call rotation can add this to the rotation as a monthly task.

The number you can defend

A backup schedule is a hope. A measured RPO is a number you can put in a runbook, compare to a target, and defend in a postmortem. The restore lottery is how you get it. Pick a minute, restore to it, and count what you lost. Then decide whether to change the configuration, change the procedure, or accept the number. All three are better than not knowing.

Most lean teams rehearse one recovery: the database. It has a dump command, a restore command, and a runbook that someone wrote after the first time it went wrong. DNS, registrar access, and console-only configuration fail the same way a database does — you discover what you lost at the moment you need it — but they have no equivalent of pg_restore, so they get skipped. This article is about treating them as first-class backup targets: inventorying what lives outside version control, capturing it in a restorable form, and rehearsing the path back.

Why the database is the easy part

The database is the easy part because the tooling is obvious. pg_dump, pg_basebackup, pgBackRest, restic with an age recipient — the ecosystem assumes you will need to restore, so it gives you a restore path. DNS and registrar access do not have that assumption baked in. A zone export tells you what the records were. It does not tell you who can publish them, which account holds the domain, or whether the registrar will let you move it today. Those are separate artifacts, and they need separate owners.

The useful framing is not “what else should I back up” but “what would I need to rebuild this service if the console tab disappeared.” For most services running on AWS, GCP, or bare metal, the answer includes at least three things that never appear in a database dump: the DNS zone, the registrar account that holds the domain, and the settings that only ever existed in a web console.

Inventory pass: what actually has no export path

Before capturing anything, list what your service depends on and mark which items have a documented export or infrastructure-as-code path. The list is usually shorter than people expect, and the gaps are usually in the same places.

  • DNS zones and records. Public zones, private zones, and any records that were created by hand during an incident.
  • Registrar account. The login, the registrant contact email, the authorization code process, and any DNSSEC DS records published at the registry.
  • Console-only configuration. Load balancer listeners and health checks, IAM policies that were edited in the console, WAF rules, CDN behaviors, queue and topic settings, and anything else that was configured by clicking rather than by committing.
  • Access. Who can change DNS, who holds the registrar credentials, and what happens to those rights when someone leaves.

The inventory is not a one-time document. It is a habit: for each resource you create, ask what you would need to rebuild it, and write that down at creation time rather than during an incident. The recovery checklist is a reasonable place to keep the output, because it is already the document you reach for when something is broken.

Capturing DNS: the export is not the recovery plan

Every major DNS provider offers some way to read your zone data. Cloud DNS publishes zones and records through its API and console, and supports IAM permissions at both the project level and the individual zone level, so read access can be granted without granting write access. Route 53 exposes hosted zones through its API and console as well. The mechanics differ, but the shape is the same: you can get the records out.

What you cannot get out is the ability to publish them. A zone export is a snapshot of record data. It does not include the registrar relationship, the nameserver delegation at the registry, or the credentials that let you change either. Treat the export as one artifact among several, not as the backup.

Two properties of DNS make the export less useful than it looks. First, changes propagate in two parts: the change must reach the authoritative name servers, and resolvers must pick it up when their cached records expire. The record TTL controls that cache. If you set a TTL of 86400, resolvers are instructed to cache for 24 hours, and some resolvers ignore TTL or use their own values. Second, many popular stub and recursive resolvers default to caching negative responses for up to 15 minutes — a behavior commonly seen in Microsoft Windows, the Java JVM, and dnsmasq. A restore that looks correct at the authoritative server can still be invisible to clients for a while. That is not a reason to skip the export; it is a reason to rehearse the restore and watch it from a client, not just from the provider’s console.

Registrar access is a credential, not infrastructure

The registrar account behaves like a credential, not like a resource. If the only person who can authorize a transfer is unreachable, the domain is effectively frozen regardless of how good your zone export is. This is the failure mode that a database-shaped backup plan does not cover, because there is no dump command for “the person who holds the login.”

What to capture, at minimum:

  • The registrar account itself, in a shared credential store rather than one engineer’s personal password manager.
  • The registrant contact email, and confirmation that it is a mailbox more than one person can read.
  • Whether the domain is locked, and who can unlock it.
  • Whether DNSSEC is enabled, and where the DS records are published.
  • The authorization code process for the current registrar, and how long it takes to obtain one.

The registrant contact matters more than it looks. AWS documents that the contact listed as registrant has certain rights as the Registered Name Holder under the ICANN Transfer Policy, and that if a domain remains in a closed AWS account, that contact might be able to request a transfer to an external registrar. The practical implication for a lean team is that the registrant contact should be a role mailbox or a trusted person, not a personal address that leaves with an employee.

What a registrar transfer actually requires

A registrar migration is the recovery scenario most teams never rehearse, and it is where the invisible constraints live. The Route 53 transfer documentation is a useful checklist even if you are not moving to Route 53, because it names the failure points that apply broadly.

The pre-transfer steps include confirming the registrant email is current, unlocking the domain, confirming the domain status allows transfer, disabling DNSSEC, obtaining an authorization code, and — for selected geographic TLDs — renewing the registration before transferring. The 60-day rule is the one that catches people: transfers are blocked within 60 days of initial registration or a registrant contact change. If you updated the registrant contact last month, you cannot move the domain this week.

DNSSEC is the second trap. For transfers to Route 53, DS records must be removed at the current registrar before transferring, and the change given at least 24 hours to propagate. If a domain registration is transferred while DNSSEC is configured and DNS service is then moved to a provider that does not support DNSSEC, resolution fails intermittently until the DNSSEC keys are deleted from the domain. That is an outage that looks like a network problem and is actually a configuration leftover.

There is also an ordering constraint worth writing into the runbook: if the registrar for your domain is also its DNS service provider, AWS recommends transferring DNS service to Route 53 or another provider before continuing with the registration transfer. Doing both at once risks taking the domain’s resolution down during the move.

And for some geographic TLDs — the Route 53 documentation lists .ch, .cl, .co.uk, .co.za, .com.au, .cz, .es, .fi, .im, .jp, .me.uk, .net.au, .org.uk, .se, and .uk — registration is not automatically extended when a domain is transferred. If the expiration date is approaching, renew before transferring, or the registration could expire mid-transfer and the domain could become available for others to purchase.

None of these constraints are visible from a zone export. A rehearsal that only checks “can I read the zone file” will not surface them. A rehearsal that walks the transfer checklist will.

Console-only configuration: a category, not a list

There is no exhaustive list of settings that lack an export button, and any article that claims one is guessing. The useful discipline is the question, not the answer: for each resource you create, what would you need to rebuild it if the console tab disappeared?

In practice, the answers cluster into a few shapes:

  • Settings that can be read but not exported. Load balancer listener rules, health check thresholds, WAF rule ordering, CDN cache behaviors. You can screenshot them or transcribe them; you cannot always get a machine-readable dump.
  • Settings that were edited in the console and drifted from code. An IAM policy that was tightened during an incident and never committed. A security group rule added by hand. These are the ones that make a rebuild-from-code produce a different system than the one that was running.
  • Settings that only exist as a relationship. Which role can assume which other role, which service account can read which secret. The individual objects may be exportable; the graph is not.

The recording method matters less than the recording. A decision record that says “the load balancer idle timeout is 120 seconds because of the long-poll endpoint, set on 2026-03-14, see ticket” is worth more during a rebuild than a screenshot, because it explains why the value is what it is. The recovery checklist can hold the pointer; the decision record holds the reasoning.

Access hygiene is backup hygiene

The offboarding checklist that removes a departing engineer’s IAM user is also the moment to confirm that the registrar contact and DNS publishing rights did not leave with them. These are the same problem viewed from two angles.

AWS recommends using IAM roles for human users and workloads so they use temporary credentials, and requiring MFA for scenarios where an IAM user or root user is needed. It also recommends updating access keys when needed — such as when an employee leaves — and using IAM access last used information to update and remove access keys safely. The same guidance applies to DNS: Cloud DNS supports IAM permissions at the project and zone level, with the DNS Administrator role (roles/dns.admin) required to make changes and the DNS Reader role (roles/dns.reader) granting read-only access. A team that grants DNS Reader broadly and DNS Administrator narrowly has a smaller blast radius when someone leaves than a team that hands out Administrator because it is simpler.

One permission detail is worth knowing before you design the break-glass path: in Cloud DNS, the DNS Administrator role does not have the setIamPolicy permission. Configuring a policy on a DNS resource such as a managed zone requires Owner access to the project that owns the resource. That means the person who can change records is not necessarily the person who can change who can change records — and the break-glass procedure needs to account for both.

AWS’s broader guidance is to regularly review and remove unused users, roles, permissions, policies, and credentials, using last accessed information to identify them. For a 2–15 engineer team, the practical version is a quarterly pass: list who has DNS Administrator, list who has registrar access, compare against who is still on the team, and remove the difference. Google Cloud’s Well-Architected Framework frames the same idea as security by design — integrating security considerations from the initial design phase rather than bolting them on — which is a reasonable description of what an access review is doing when it is done on a schedule instead of after an incident.

Rehearsal: a restore lottery for DNS and registrar access

A restore lottery for DNS should be scored on the same terms as a database restore lottery: did someone who did not build the system reconstruct a working zone from the artifacts alone, within a time budget, without asking the person who set it up?

What “passing” looks like, concretely:

  • The person running the drill can find the zone export without asking where it is.
  • They can identify the registrar account and the registrant contact from the artifacts, not from memory.
  • They can state whether DNSSEC is enabled and where the DS records live.
  • They can walk the transfer checklist and name the constraints that would block a transfer today — the 60-day window, an unlocked domain, a current authorization code.
  • They can resolve a name against the restored zone from a client, not just from the provider’s console, and they understand that negative caching may delay what they see.

What “failing” looks like is more instructive. A team that stores its zone export in the same repository as its application code, but keeps the registrar login in a password manager entry owned by one engineer’s personal account, has a backup that survives the engineer’s departure and an access model that does not. The failure is not in the backup; it is in the access model around the backup. A team that rehearses a registrar transfer and discovers mid-rehearsal that DNSSEC is enabled and the DS records were never documented has found the problem in the cheap place. The incident is the expensive one.

Decision record: what you chose not to back up

The last artifact is the one that says what you deliberately left out. Not every console setting is worth capturing. A team that tries to back up everything will produce a document nobody reads; a team that backs up nothing outside the database will discover the gap during an incident. The decision record is where you write down which side of that line you chose and why.

A useful decision record for this topic answers four questions:

  1. Which DNS zones and records are in scope, and which are considered disposable?
  2. Who holds registrar access, and what is the break-glass path if they are unreachable?
  3. Which console-only settings are documented, and which are considered reconstructible from code?
  4. When will this be revisited — after the next offboarding, after the next provider change, or on a fixed cadence?

The revisit trigger matters more than the cadence. A registrar change, a DNSSEC enablement, or a departure from the team are all events that invalidate assumptions in the record. Writing down the trigger is what keeps the record from becoming a document that was true once.

FAQ

Is a zone export enough to recover DNS?

No. The export captures record data. It does not capture the registrar relationship, the nameserver delegation at the registry, the credentials that let you publish changes, or whether DNSSEC is enabled. Those are separate artifacts with separate owners.

How often should we rehearse a DNS or registrar recovery?

There is no sourced number, and any specific interval would be a guess. The useful trigger is change: after a registrar migration, after enabling or disabling DNSSEC, after a change to who holds registrar access, and after any offboarding that touches DNS permissions. A fixed cadence is a reasonable backstop, but the event-driven triggers are what catch the drift.

Can we back up console-only settings with an API?

Sometimes. The retrieved provider documentation does not enumerate which console-created resources have an export or infrastructure-as-code path, so the honest answer is that it depends on the resource. The discipline that works regardless is to ask, at creation time, what you would need to rebuild the resource, and to record the answer where the recovery checklist can find it.

What is the most common failure point in a registrar transfer?

The Route 53 documentation names several: email not received, domain locked, invalid authorization code, and the 60-day rule. The 60-day rule is the one that is invisible until you hit it, because it is triggered by events — initial registration or a registrant contact change — that may have happened months earlier and been forgotten.

Does DNSSEC affect recovery?

Yes, in two ways. For transfers to Route 53, DS records must be removed at the current registrar before transferring, and the change given at least 24 hours to propagate. And if a domain registration is transferred while DNSSEC is configured and DNS service is then moved to a provider that does not support DNSSEC, resolution fails intermittently until the DNSSEC keys are deleted from the domain. Both are configuration leftovers that look like network problems.

Who should hold registrar access on a small team?

The retrieved sources do not prescribe a team size or a role. The constraint that matters is that the registrant contact should be a mailbox or a person who will still be reachable when the domain needs to move — not a personal address that leaves with an employee. AWS documents that the registrant contact has rights as the Registered Name Holder under the ICANN Transfer Policy, which is why the choice of contact is a recovery decision, not an administrative one.

On a team of two to fifteen engineers, the person who pushed the change, the person who answered the page, and the person who writes the postmortem are usually the same person. That is not a moral problem. It is a documentation problem with a known failure mode: the postmortem either becomes a confession, or it does not get written at all.

The sources are clear about what a postmortem is for. Google’s SRE book defines it as “a written record of an incident, its impact, the actions taken to mitigate or resolve it, the root cause(s), and the follow-up actions to prevent the incident from recurring.” It also states that “writing a postmortem is not punishment—it is a learning opportunity for the entire company.” Those two sentences sit in tension when you are the author of the change. The format names the actions that led to the incident; the culture forbids indicting the person who took them.

That tension is not a reason to soften the record. It is a reason to be more precise about where the causal chain actually runs.

Blameless does not mean actionless

The SRE book is explicit: “For a postmortem to be truly blameless, it must focus on identifying the contributing causes of the incident without indicting any individual or team for bad or inappropriate behavior. A blamelessly written postmortem assumes that everyone involved in an incident had good intentions and did the right thing with the information they had.”

Read that carefully. Blameless does not mean the actions disappear. The same chapter acknowledges that “blameless postmortems can be challenging to write, because the postmortem format clearly identifies the actions that led to the incident.” The actions stay in the document. What changes is the question you ask about them.

The question is not “why did I do that.” The question is “what did the system show me at that moment, and what would have made the correct action the default.”

This matters more for a self-authored postmortem, not less. The SRE workbook’s case study of a bad postmortem is blunt about the cost of individual highlighting: “It may seem like a good idea to highlight individuals in a postmortem. Instead, this practice leads team members to become risk-averse because they’re afraid of being publicly shamed. They may be motivated to cover up facts critical to understanding and preventing recurrence.” On a small team, the person most likely to cover up a fact is the person who wrote the deploy. If the postmortem reads as a confession, the next one does not get written.

A three-column causal chain

The practical move is to write the causal chain in three columns before you write prose. This is a working method, not a source-prescribed template, but it follows the sources’ emphasis on contributing causes over individual fault.

Column one: what the operator did. State it plainly. “Ran the migration command a second time after the first attempt returned a timeout.” No adjectives. No “carelessly” or “stupidly.” The workbook’s bad-postmortem example flags “careless ignorance” as superfluous language that distracts from the key message.

Column two: what the system showed at that moment. This is the column most self-authored postmortems skip. What did the terminal output say? Did the timeout message distinguish between “the command did not run” and “the command ran but did not return”? Was the runbook step ambiguous about whether a retry was safe? Did the deploy tool have a lock, a confirmation prompt, or an idempotency key? If the answer is “the system showed nothing that would have stopped me,” write that down. It is a finding.

Column three: what change would have made the correct action the default. This is where action items come from. The workbook states the principle directly: “In general, trying to change human behavior is less reliable than changing automated systems and processes.” An action item that says “be more careful with migrations” is not an action item. An action item that says “add a preflight check to the migration script that refuses to run if the target table already has a lock row from a prior attempt” is one.

Only column three produces work that survives a personnel change. Columns one and two are the record; column three is the fix.

The template sections, in order

The SRE book’s definition gives the required sections: incident record, impact, actions taken to mitigate or resolve, root cause(s), and follow-up actions to prevent recurrence. The workbook’s case study adds specific failure modes to avoid in each.

Incident record and context

Include a background or glossary section if your service uses internal names that a new hire would not know. The workbook warns: “If you don’t properly contextualize content when writing a postmortem, the document might be misunderstood or even ignored. It’s important to remember that your audience extends beyond the immediate team.” On a small team, the audience beyond the immediate team is the person you hire next quarter.

Impact

Put numbers in it. The workbook is direct: “For outages affecting multiple services, you should present numbers to give a consistent representation of impact… Even if there is no concrete data, a well-informed estimate is better than no data at all.” For a lean team, that might be “approximately 40 minutes of elevated 5xx rate on the checkout endpoint, affecting an estimated 3% of requests during the window.” The estimate is better than the absence.

Actions taken to mitigate or resolve

Do not leave this empty. The workbook’s bad-postmortem example left the recovery efforts section blank and flagged it as a key omission. Write what you actually did: rolled back the deploy, restarted the worker pool, failed over the read replica. This section is also where you record what worked. The workbook notes that a good postmortem lets readers know “what happened, how the issue was mitigated, and how users were impacted.”

Root causes and trigger

This is where the three-column method does its work. The trigger is the event that started the incident. The contributing causes are the conditions that let the trigger produce the impact. The workbook’s case study criticizes a postmortem that “contains a small paragraph that describes the root causes and trigger, but it doesn’t explore the lower-level details of the issue.” Go lower. If the root cause is “the migration script was not idempotent,” the contributing causes include why the script was written that way, why no test caught it, and why the runbook did not warn against a retry.

Action items

The workbook lists the failure modes precisely. Action items should include preventative items, not only mitigative ones. They should use concrete verbs instead of “improve” or “make better,” because “these terms are vague and open to interpretation” and “make it difficult to measure and understand success criteria.” They should have differentiated priorities, not all P2. They should have a tracking bug, because “without a formal tracking process, action items from postmortems are often forgotten, resulting in outages.” And they should have a single owner: “Ideally, an owner is a single point of contact who is responsible for the postmortem, follow-up, and completion.”

For a team of two to fifteen, the single owner is usually the person who wrote the postmortem. That is fine. The point is that the name is on the item, not that the name is on the incident.

Review when you have no review committee

The SRE book states: “An unreviewed postmortem might as well never have existed.” It describes a review process where teams “share the first postmortem draft internally and solicit a group of senior engineers to assess the draft for completeness,” with criteria including whether key incident data was collected, whether impact assessments are complete, whether the root cause is sufficiently deep, whether the action plan is appropriate, and whether the outcome was shared with relevant stakeholders.

A team of five does not have a group of senior engineers to convene. The lean adaptation, which is a proposal rather than a source finding, is a single named reviewer who was not involved in the incident, plus a short written checklist drawn from those criteria. The reviewer’s job is not to judge the author. It is to ask whether the root cause section goes deep enough and whether the action items have owners and tracking bugs. One reviewer, one pass, one checklist.

Publish promptly. The workbook’s bad-postmortem example was published four months after the incident, and “in the interim, had the incident recurred (which in reality, did happen), team members likely would have forgotten key details that a timely postmortem would have captured.” On a small team, the details you lose in four months are the ones you need most: the exact command, the exact error message, the exact state of the system.

Share as widely as is useful. The workbook states: “The value of a postmortem is proportional to the learning it creates. The more people that can learn from past incidents, the less likely they are to be repeated.” For a lean team, that might mean the whole engineering group, or it might mean the two people who will be on call next month. The point is that the document is not private to the author.

What this does not fix

A blameless postmortem does not make the incident less annoying. It does not change the fact that you were the one who ran the command. It does not, on its own, prevent the next incident. The SRE book describes blameless postmortems as “a tenet of SRE culture” and the workbook describes them as important to “creating and maintaining a successful SRE organization,” but neither source claims that writing one is sufficient. The PagerDuty postmortem guide frames the practice as one that “allow[s] your teams to iteratively improve your infrastructure and incident response process” — iteratively, not instantly.

What it does is keep the record honest. The alternative, on a small team, is a postmortem that either blames the author or omits the author’s actions. The first teaches the team that incidents are personal. The second teaches them nothing. The SRE book’s warning is the one to keep in view: “An atmosphere of blame risks creating a culture in which incidents and issues are swept under the rug, leading to greater risk for the organization.”

Write the actions down. Put the system in the causal chain. Give the action items owners and tracking bugs. Then publish it before you forget the error message.

FAQ

Does a blameless postmortem mean I should not name what I did?

No. The SRE book states that the postmortem format “clearly identifies the actions that led to the incident.” The actions stay. What changes is that the postmortem focuses on contributing causes without indicting the individual, and assumes the person acted with good intentions and the information they had.

What if the only honest root cause is that I ran the command again?

Write that sentence, then write the next one: why was running it again possible, and what would have stopped it? The workbook’s principle is that changing automated systems and processes is more reliable than trying to change human behavior. The second sentence is where the action item lives.

How many action items should a postmortem have?

The sources do not prescribe a number. The workbook’s criteria are about quality: at least one preventative item, concrete verbs, differentiated priorities, a tracking bug, and a single owner per item. A postmortem with three well-owned items is more useful than one with ten vague ones.

Who should review a postmortem on a team with no SRE department?

The sources describe review by senior engineers against a set of completeness criteria. A lean-team substitute, which is a proposal rather than a source finding, is one named reviewer who was not involved in the incident, working from a short checklist: was the data collected, is the impact assessed, is the root cause deep enough, are the action items owned and tracked, and were the relevant people told.

Should I write a postmortem for a near miss?

The SRE book lists common postmortem triggers including user-visible downtime, data loss, on-call intervention, resolution time above a threshold, and monitoring failure. It also notes that “any stakeholder may request a postmortem for an event.” A near miss that would have met one of those triggers if a single condition had differed is a reasonable candidate, but the sources do not prescribe a rule for near misses specifically.

A new service is about to take its first real traffic. For a team of two to fifteen engineers who are also the on-call rotation, the first incident on that service is usually a cold start: nobody has decided who gets paged, what they look at first, or which recovery path they attempt. A postmortem after that incident will produce useful learning, but the learning is paid for in user impact. A pre-mortem is the prospective mirror of that practice — same blameless ground rules, same insistence on owned and tracked action items, same one-page record, applied before the service is exposed rather than after it fails.

This is a proposed adaptation, not a documented industry practice. The postmortem craft it borrows from is well documented; the pre-mortem format below is a recommendation for lean teams that lack dedicated reliability coverage.

Why a pre-mortem, not another launch checklist

PagerDuty’s postmortem guide frames the after-the-fact practice plainly: performing postmortems after incidents is how teams learn what they are doing right, where they could improve, and how to avoid repeating mistakes. The same guide notes that a successful postmortem process rests on a culture of honesty, learning, and accountability, and that culture change can be led from any role even though it benefits from management buy-in.

A pre-mortem applies that same posture to a service that has not yet failed. The Google SRE book states that for a postmortem to be truly blameless it must focus on identifying contributing causes without indicting any individual or team, and that a blamelessly written postmortem assumes everyone involved had good intentions and did the right thing with the information they had. Those two sentences are the entire ground rule for the pre-mortem. The team is not predicting who will break the service; it is enumerating the conditions under which the service, as designed, will fail, and deciding in advance what the response will be.

The Google SRE book also notes that it is important to define postmortem criteria before an incident occurs so that everyone knows when a postmortem is necessary. The pre-mortem extends that logic one step earlier: define the failure modes and the response criteria before the service takes traffic, so the first incident is a rehearsal of a known plan rather than an improvised one.

The 45-minute bound is a design constraint

Forty-five minutes is not a target to fill. It is a constraint that forces the team to cluster and rank failure modes instead of enumerating every conceivable one. A session that runs ninety minutes will be scheduled once and then skipped. A session that runs forty-five minutes can be run before every new service, not only before the largest launches.

The agenda below assumes four to eight participants. One person facilitates and one person writes. The facilitator does not need to be the most senior engineer; the writer does not need to be the service owner.

Minutes 0–5: framing and scope

State the service, the traffic it is about to receive, the data it will touch, and the people who will be paged when it fails. Name the blast radius explicitly: which other services depend on this one, and which users notice. Write the scope on the shared document so it does not drift.

Minutes 5–15: silent failure-mode generation

Everyone writes failure modes independently, in silence, into the shared document. Silent generation prevents the first confident voice from anchoring the room. The prompts below are the ones that fit lean teams; use them as starting points, not as a fixed list.

  • What breaks at 03:00 when only one person is awake?
  • What happens if the person who built the service is on holiday?
  • What does the first bad deploy look like?
  • What does the first bad migration look like?
  • What does the first credential rotation look like?
  • What does the first cold cache look like under real traffic?
  • What does the first connection-pool exhaustion look like?
  • What does the first missing index look like on the largest table?
  • What does the first rate-limit or auth surprise look like at the edge?

Minutes 15–30: clustering and ranking

Group the failure modes into clusters. Rank each cluster by likelihood and blast radius. The ranking does not need to be precise; it needs to be explicit enough that the team agrees on which three or four clusters deserve action items this week and which can be deferred. A cluster that survives ranking should produce at least one action item.

Minutes 30–40: owners and tracking

Each action item gets a single owner, a priority, and a tracking artifact. The Google SRE workbook states that action items without clear owners are less likely to be resolved, and that it is better to have a single owner and multiple collaborators. The same workbook states that without a formal tracking process, action items from postmortems are often forgotten, resulting in outages. The pre-mortem inherits both rules. Vague verbs like “improve” or “make better” are not action items; they are intentions. An action item names a specific change, a specific owner, and a specific place where its completion is visible.

The workbook also states that trying to change human behavior is generally less reliable than changing automated systems and processes. When a failure mode resolves to “the on-call engineer should remember to do X,” the action item should be rewritten to make X automatic, or to make the omission visible in monitoring, rather than to ask a tired human to remember.

Minutes 40–45: the one-page record

Write the record while the session is still open. The record contains: scope, participants, failure modes considered, action items with owners and tracking, and a date to revisit. The Google SRE workbook states that a postmortem is a factual artifact that should be free from personal judgments and subjective language, should consider multiple perspectives, and should be respectful of others. The pre-mortem record follows the same standard. It is a document, not a transcript of who said what.

Turning imagined failures into rehearsals

An action item that only changes a document is cheap and often sufficient. An action item that changes what the team has actually done before launch is stronger. For each failure mode that survives ranking, ask two questions: which signal would reveal it, and which rehearsal would exercise the recovery path?

Consider a hypothetical team about to expose a new service to real traffic. One failure mode they might imagine is that the only engineer who knows how to restore the primary database is unreachable at 03:00. The action item is not a note in a wiki; it is a restore lottery scheduled before launch, in which a different engineer performs the restore from a recent snapshot on a scratch host. The restic documentation describes the restore command and the use of the word latest to select the most recent snapshot, and it warns that restoring data in-place can leave files in a partially restored state if the operation is interrupted, recommending a current backup before restoring a different snapshot. A restore lottery that runs against a scratch host, not against production, is the rehearsal that turns that warning into a known procedure.

Another failure mode a team might imagine is that the first bad migration leaves the database in a state that requires point-in-time recovery. The PostgreSQL documentation states that PostgreSQL maintains a write ahead log in the pg_wal/ subdirectory of the cluster’s data directory, and that the log records every change made to the database’s data files. It states that continuous archiving supports point-in-time recovery, making it possible to restore the database to its state at any time since the base backup was taken. It also states that to recover successfully using continuous archiving, you need a continuous sequence of archived WAL files that extends back at least as far as the start time of your backup, and that you should set up and test your procedure for archiving WAL files before you take your first base backup. The pre-mortem action item is therefore concrete: verify before launch that WAL archiving is running, that a base backup exists, and that at least one engineer has performed a PITR restore on a non-production copy.

The same documentation notes that if archiving falls significantly behind, the amount of data that would be lost in a disaster increases, and the pg_wal/ directory will contain large numbers of not-yet-archived segment files that could eventually exceed available disk space. It also notes that the archive command is only invoked on completed WAL segments, and that archive_timeout can force the server to switch to a new segment at least that often. A pre-mortem that names “WAL archiving silently falls behind” as a failure mode should produce an alert on archive lag and a documented archive_timeout value, not a reminder to check the directory.

A third hypothetical failure mode: the first credential rotation locks out the deploy pipeline. The action item is a one-page runbook that names the break-glass path and the person who holds the second key. The runbook is written before launch, not during the incident. The Recovery Checklist Before You Need It is a useful companion here, because it forces the team to write the recovery steps down while the service is still healthy.

What the pre-mortem is not

The pre-mortem is not a substitute for the postmortem. The Google SRE workbook states that when written well, acted upon, and widely shared, postmortems can be a very effective tool for driving positive organizational change and preventing repeat outages. The pre-mortem reduces the number of surprises; the postmortem after the first real incident remains the place where the team learns what the rehearsal missed. The two practices share a repository, a tone, and a standard for action items, but they answer different questions.

The pre-mortem is also not a launch coordination checklist, a canary release, or a rollback plan. It does not replace any of those. It sits alongside them and produces the action items that make them specific to this service.

Nor is the pre-mortem a prediction. The team is not trying to name the exact failure that will occur. It is pre-deciding who will be paged, what they will look at first, and which recovery path they will attempt. The value is in the pre-decision, not in the accuracy of the forecast.

Closing the loop

Schedule the review before the session ends. The Google SRE workbook states that the value of a postmortem is proportional to the learning it creates, and that the more people who can learn from past incidents, the less likely they are to be repeated. The pre-mortem record should be shared as widely as the team’s postmortems are shared, and it should live in the same repository. When the first real incident occurs, the team compares what happened against what was anticipated. The gap between the two is the input to the next pre-mortem.

For a team that is its own on-call rotation, the pre-mortem is the cheapest rehearsal available. It costs forty-five minutes and a one-page document. It does not require a dedicated SRE function, a new tool, or a vendor. It requires only that the team treat the first real traffic as a rehearsal of a known plan rather than a cold start.

Questions readers ask

How is a pre-mortem different from a design review?

A design review evaluates whether the service is built correctly. A pre-mortem assumes the service is built as designed and asks how it will fail under real traffic, who will be paged, and what the response will be. The two meetings can be adjacent, but they produce different artifacts.

Do we need a facilitator who is not on the team?

No. The facilitator’s job is to keep the timebox and to prevent the first confident voice from anchoring the room. Any team member can do it. The writer’s job is to keep the record factual and free of personal judgment, which is the same standard the Google SRE workbook applies to postmortems.

What if we cannot rank failure modes by likelihood?

Rank by blast radius instead. The point of ranking is to decide which clusters get action items this week. A cluster that would take down the service for every user outranks a cluster that would degrade one internal dashboard, even if the likelihood is unknown.

How many action items should come out of a 45-minute session?

Enough to cover the clusters that survived ranking, and no more. The Google SRE workbook states that action items without clear owners are less likely to be resolved. A short list with single owners and tracking artifacts is more useful than a long list of intentions.

Should the pre-mortem record be public?

Share it as widely as the team’s postmortems are shared. The Google SRE workbook states that the value of a postmortem is proportional to the learning it creates. The same logic applies to the pre-mortem record, with the caveat that it may contain details about unreleased services that should not leave the organization.

Nobody scribed. That is the normal case, not the failure case. The incident channel has 400 messages, the alert history has a dozen state transitions, and the application logs have a few thousand lines. The postmortem is due in three days. What you have is not a memory problem; it is a reconstruction problem, and reconstruction is a procedure.

The useful framing is that you are reconciling three independent clocks. PagerDuty records when alerts fired, were acknowledged, and resolved. Your log platform records when the system did things. Slack records when people said things. Only two of those three are reliable for ordering events, and the third is the only source for intent. Treating them as interchangeable is what produces timelines that look precise and are wrong.

Name an owner before the channel goes quiet

The single highest-leverage move for a 2–15 engineer team is not a new tool. It is designating a postmortem owner at incident close, while the responders are still in the channel. PagerDuty’s public postmortem process describes this role explicitly: the owner is responsible for populating the postmortem, looking up logs, managing the follow-up investigation, and keeping interested parties in the loop (PagerDuty, Postmortem Process).

That sentence converts “nobody took notes” from a memory failure into a scheduled task. The incident commander can name the owner in the final minutes of the call. The owner does not need to have been the deepest responder; they need to be the person who will spend the next afternoon in query consoles and Slack search.

Google’s SRE book frames the same artifact more broadly: a postmortem is a written record of an incident, its impact, the actions taken to mitigate or resolve it, the root cause(s), and the follow-up actions to prevent recurrence (Google SRE Book, Chapter 15). The timeline is the spine that holds those sections together. Without it, the analysis floats free of the events.

Set the reconstruction window and pick the authoritative clock

Before opening any tool, write down the window. Start it earlier than the first alert — the causal chain usually begins before the first page — and end it after the last state transition, not after the last Slack message. A window that starts at the first alert and ends at the last chat message will systematically miss the change that caused the incident and the cleanup that followed it.

Then decide which clock governs which class of event:

  • Alert and escalation transitions: the monitoring platform’s timestamps. These are machine-generated and monotonic within the platform.
  • System behavior: log timestamps, with the caveat that ingestion delay and clock skew are real. Note the offset between the log platform’s receive time and the event’s own timestamp field where both exist.
  • Human decisions: Slack message times, treated as approximate. They tell you ordering within a conversation, not absolute time.
  • Database recovery events: PostgreSQL Log Sequence Numbers, which increase monotonically with each WAL record and can be compared to measure the volume of WAL data between two points (PostgreSQL Documentation, WAL Internals). LSNs order database events precisely. They say nothing about human decisions, and you should not ask them to.

Writing this mapping down once, in the incident channel or the postmortem draft, prevents the most common reconstruction error: quoting a Slack timestamp as if it were an event time.

Pull the machine record first

Query the log platform before you read a single chat message. The machine record gives you a monotonic spine, and reading Slack first anchors your reconstruction on the least reliable clock.

If you are on CloudWatch Logs Insights, the query language supports the operations you need for this pass. The filter command returns only log events matching one or more conditions. The sort command orders results ascending or descending. The limit command caps the number of events returned, which is useful with sort to get the most recent or top results rather than the full window (AWS, CloudWatch Logs Insights query syntax).

Two commands are worth knowing for reconstruction specifically. join combines log events from a source log group with events from another log group or query result based on a matching field, which is how you correlate related events across sources using a request identifier or transaction ID. sessionize groups events into sessions by identity fields and an inactivity gap, which is how you turn a stream of retries into a small number of logical attempts. Both are documented in the same reference.

If the incident involved a database recovery, pull the backup tool’s logs next. pgBackRest’s user guide notes that a restore requires the backup files and one or more WAL segments to work correctly, and that WAL segments follow a naming convention where the first eight hexadecimal digits represent the timeline and the next sixteen are the logical sequence number (pgBackRest User Guide). Those names are themselves a timeline. If you restored from a restic repository, the troubleshooting section of the restic documentation covers finding damaged data, backing up the repository, repairing the index, and checking the repository again (restic Documentation) — useful when the reconstruction reveals that the restore itself had a gap.

One operational note worth carrying into the timeline: pgBackRest requires that local and remote versions match exactly, and a mismatch stops WAL archiving and backups from functioning until versions match. If your incident involved a failed archive, the version mismatch is a candidate cause and belongs in the timeline as an observed fact, not a hypothesis.

Mine Slack second, and read it narrowly

Slack is the only source for intent. It tells you who decided what, when a hypothesis was abandoned, and which workaround was tried before the one that worked. It is a poor source for what happened when.

PagerDuty’s process is direct about this use: go through the history in Slack to identify the responders and add them to the page. Identify the incident commander and scribe in that list. The same process treats the timeline as the main focus when beginning to populate the postmortem, and specifies that it should include important changes in status or impact and key actions taken by responders.

Read Slack with a purpose. You are looking for four things:

  1. Decision points. “Rolling back the deploy” is a timeline entry. The discussion that preceded it is context.
  2. Abandoned hypotheses. These are the most valuable and most commonly lost entries. A hypothesis that was tested and rejected is a fact about the incident, and it prevents the next responder from re-testing it.
  3. Presence. Who joined the channel, when, and from which team. This is the map of who can corroborate which rows.
  4. Workarounds. What was tried, in what order, and what the observed effect was.

Do not read Slack for timestamps. Read it for content, then attach the nearest machine timestamp as the row’s time and mark the row as approximate.

Reconcile into one table with a source column

The output of the reconstruction is a table, not prose. Each row has a time, an event, a source, and a flag for whether the row is observed or inferred.

A minimal schema:

Time Event Source Observed / Inferred
T+00:00 First alert fires PagerDuty alert history Observed
T+00:14 Escalation to secondary PagerDuty escalation history Observed
~T+00:22 Failover initiated Slack #inc-1234, time approximate; corroborated by PagerDuty escalation at T+00:14 Inferred
T+00:31 Error rate returns to baseline CloudWatch Logs Insights query, permalink in appendix Observed

The third row is the important one. A hypothetical example, clearly labeled as such: a team discovers at postmortem time that the only record of a failover decision is a one-line Slack message with no timestamp context. Under this approach the row reads as above — approximate time, named source, corroborating machine event — rather than a confident but unsourced time. A visible gap is more useful to a reviewer than a plausible sentence.

PagerDuty’s process asks for exactly this discipline: for each item in the timeline, identify a metric or some third-party page where the data came from — a graph link, a search, a post — anything that shows the data point you are trying to illustrate. It also asks that any commands or queries used to look up data be posted on the page so others can see how the data was gathered. That second requirement is what makes the timeline reviewable rather than merely readable.

Write each row so a reviewer can re-derive it

A timeline entry that cannot be re-derived is an assertion. The test is simple: can a reviewer who was not on the call take the row’s source, run the query or open the permalink, and see the same thing?

For log-derived rows, paste the query. For alert-derived rows, link the incident in the monitoring platform. For Slack-derived rows, link the message permalink. For database recovery rows, cite the LSN range or the WAL segment names. The cost of this discipline is a few minutes per row. The benefit is that the postmortem meeting does not spend its first fifteen minutes arguing about whether an event happened at T+22 or T+31.

Google’s SRE book is blunt about the alternative: an unreviewed postmortem might as well never have existed. The review is what converts a document into a shared understanding, and a timeline without sources cannot be reviewed — only believed or disbelieved.

Circulate the draft before the meeting

PagerDuty’s process recommends posting a link to the postmortem into Slack for internal review of style and content roughly 24 hours before the meeting, so that experienced readers can flag missing detail while there is still time to fix it. For a small team, this is the difference between a meeting that resolves disagreements and a meeting that discovers them.

The review pass has a specific job for the timeline: check that every row’s source actually supports the row. This is where inferred rows get promoted to observed, or get marked as gaps. It is also where a reviewer who was on the call can supply a missing corroborating source — a graph they had open, a query they ran — that the owner could not have known about.

The meeting itself, per the same process, opens by recapping the timeline to make sure everyone agrees and is on the same page. That recap is short when the timeline is sourced and long when it is not.

Keep the recipe as a runbook

The reconstruction procedure — the window definition, the clock mapping, the query patterns, the table schema, the review window — is itself a runbook artifact. Written down once, the next incident’s timeline costs an afternoon instead of a week of intermittent effort. It can also be rehearsed: a pre-mortem that walks through the reconstruction steps on a past incident surfaces the gaps in your logging and your chat conventions before you need them under pressure.

This connects to a broader practice worth adopting: writing the recovery checklist before you need it. The same logic applies to postmortem reconstruction. The checklist you write during a calm week is the one you can follow during a bad one.

Two conventions make the recipe durable. First, keep the timeline table in the postmortem template, with the source and observed/inferred columns already present, so the owner does not have to invent the schema under time pressure. Second, keep a short appendix of the queries used, with the incident window parameterized, so the next owner can adapt rather than rediscover.

What this does not fix

Reconstruction is not a substitute for contemporaneous notes. A scribe who captures decisions in real time produces a better timeline than any amount of after-the-fact querying. But a small team without a dedicated SRE function will not always have a scribe, and the reconstruction procedure is what keeps that from becoming a postmortem that never gets written.

It also does not fix missing data. If the log retention window closed before the postmortem owner started, the row is a gap. Label it as a gap and move on. The gap itself is a finding — it tells you something about retention policy that is worth a follow-up ticket.

Finally, the procedure does not make the timeline blameless by itself. Blamelessness is a property of how the rows are written and how the meeting is run. Google’s SRE book defines a blameless postmortem as one that focuses on identifying the contributing causes of the incident without indicting any individual or team for bad or inappropriate behavior. A timeline that records “engineer X ran the wrong command” is not blameless; one that records “the runbook’s rollback step assumed a different deploy mechanism than the one in use” is.

FAQ

What if the incident channel was archived or the messages are gone?

Then Slack is not a source for that incident, and the timeline is built from the machine record alone. Mark the human-decision rows as gaps. The reconstruction procedure still works; it just produces a thinner timeline. The follow-up action is a retention or export decision, not a reconstruction technique.

How do I handle timezone differences between responders?

Normalize every timestamp to UTC in the table, and note the original timezone in the source column if it matters. The machine sources — alert history, log timestamps — are typically already UTC or carry an explicit offset. Slack displays in the viewer’s local time, which is one more reason to treat its timestamps as approximate.

Should the timeline include events that turned out to be irrelevant?

Yes, if they consumed responder attention. An abandoned hypothesis is a fact about the incident. The timeline’s job is to record what happened, including the dead ends, so the next responder does not re-walk them.

How detailed should each row be?

Detailed enough that a reviewer can re-derive it from the cited source. If the source is a query, the row needs the query. If the source is a graph, the row needs the link and the time range. If the source is a Slack message, the row needs the permalink and a note that the time is approximate.

What if the postmortem owner was also a responder?

That is common on small teams and it is workable, with one caveat: the owner’s own actions are the hardest rows to source, because the owner remembers them and may not have written them down. Flag those rows explicitly and ask a second responder to corroborate them during the review pass.

Does this work for a database recovery incident specifically?

Yes, and the database gives you better anchors than most incidents. PostgreSQL LSNs are monotonic and comparable, so you can order recovery events precisely. pgBackRest WAL segment names encode the timeline and logical sequence number, so the archive history is itself a timeline. The human decisions around the recovery — when to fail over, when to accept data loss — still come from Slack and still need the approximate-time treatment.

The short version

Name a postmortem owner at incident close. Define the window and the authoritative clock for each event class. Pull the machine record first, Slack second. Reconcile into a table with a source column and an observed/inferred flag. Cite a query, graph, or permalink for every row. Circulate the draft a day before the meeting. Keep the recipe as a runbook. The timeline that results is not a reconstruction of memory; it is a reconstruction of evidence, and it is reviewable by someone who was not there.

On a five-engineer rotation, every page lands on the same small pool of people. That is the constraint that makes severity definitions matter more for you than for a company with a dedicated incident-response org. A big-company SEV ladder assumes you can staff an incident commander, a communications lead, and an ops lead for every major incident. You cannot. So your SEV table has to do double duty: it decides how hard you respond, and it decides whether the on-call engineer wakes anyone else up.

The common failure mode is defining severity by who got paged rather than by measured impact. When one customer’s export job fails and the whole API is down, both arrive on the same phone. If your definitions do not distinguish them, the on-call engineer is left to improvise at 3 a.m. — which is exactly when improvisation is most expensive.

Severity is impact. Priority is urgency.

Atlassian draws a distinction that is worth borrowing even if you never use their tooling: severity is a measurement of impact, and priority is a measurement of urgency. A typo on the homepage is low severity and possibly high priority. An app crash affecting 0.05% of users is high severity and possibly not the top priority if something wider is also burning. The two fields answer different questions, and conflating them is how teams end up arguing about labels during an incident instead of resolving it.

For a small team, the practical consequence is that your SEV level should be a statement about measured impact — how many customers, which functionality, how much redundancy is gone — not a statement about how annoyed the on-call engineer is. Priority is the field you use to sequence work. Severity is the field you use to decide response intensity.

What a metric-driven ladder looks like

PagerDuty’s public incident-response documentation is a useful reference because it is specific. Their SEV-1 covers a critical issue warranting public notification and executive liaison, with functionality severely impaired for a large number of customers or a customer-data-exposing vulnerability. SEV-2 covers a critical system issue actively impacting many customers’ ability to use the product. Anything above SEV-3 is automatically considered a major incident and gets a more intensive response. SEV-3 covers partial loss of functionality not affecting the majority of customers, or something with the likelihood of becoming a SEV-2 if nothing is done. SEV-4 covers performance issues, individual host failure, delayed job failure, and cron failure that does not impact the event and notification pipeline. SEV-5 is cosmetic.

Two things in that ladder are worth copying directly. First, the definitions are metric-driven. PagerDuty explicitly recommends making your own definitions very specific, usually referring to a percentage of users or accounts affected. Second, the tie-breaker: if you are unsure whether an incident is SEV-2 or SEV-1, treat it as the higher one. During an incident is not the time to litigate severities; review the classification during the postmortem.

Atlassian’s ladder is similar in shape. Their SEV-1 examples include a customer-facing service down for all customers, a confidentiality or privacy breach, and customer data loss. SEV-2 includes a customer-facing service unavailable for a subset of customers and core functionality significantly impacted. SEV-3 includes a minor inconvenience with a workaround available and usable performance degradation. Atlassian also notes that SEV-3 incidents can be handled during working hours, while SEV-1 and SEV-2 generate an alert for on-call professionals regardless of time of day.

Adapting the ladder for five people

The mistake is to copy a five-tier ladder and assume it fits. It may not. Atlassian’s own guidance is that when setting severity levels you need to factor in the size of your tech team, your on-call schedules, your high- and low-traffic times, and the frequency of incidents. A five-person team with a single time zone and a service that is quiet between 2 a.m. and 7 a.m. has different constraints than a global service with follow-the-sun coverage.

A workable approach for a small team is to define three or four levels, each tied to a measurable condition, and to decide explicitly which levels page. A starting shape, adapted from the sources above:

  • SEV-1: Core functionality is unavailable for all or nearly all customers, or customer data is lost or exposed. Pages the on-call engineer and wakes a second person. Public or customer-facing communication is warranted.
  • SEV-2: Core functionality is significantly impaired for a subset of customers, or a critical internal pipeline is broken. Pages the on-call engineer. Escalation to a second person is at the on-call engineer’s discretion.
  • SEV-3: Partial loss of functionality not affecting the majority of customers, or a condition that will become SEV-2 if nothing is done. Pages during business hours; may page out of hours if the on-call engineer judges it necessary.
  • SEV-4: Performance degradation, single host failure, delayed job failure, or cron failure with no customer-visible impact. Ticket, not page.

The percentages and the specific functionality names are yours to fill in. The point is that each level has a written condition that a tired engineer can check against reality without a meeting.

The single-customer outage

This is the case the topic names, and it is where a metric-driven definition earns its keep. Suppose your service has 200 customers and one reports that a report export is failing. That is a partial loss of functionality for a small subset. Under a metric-driven definition it is a SEV-3 or SEV-4, not a SEV-1 — even though it arrived on the same pager as an everyone-down incident would.

The useful question is not “is this a SEV-1?” It is “does this meet our written threshold for a major incident?” If the answer is no, it can be handled as a SEV-3 or SEV-4 during business hours without waking the whole rotation. If the answer is yes — because the export is core functionality for a segment that represents a meaningful share of revenue, or because the failure is a symptom of something wider — then it escalates on the merits, not on the volume of the complaint.

This is also where the “assume the worst” tie-breaker helps a small team specifically. Debating severity during an incident has a social cost: someone has to argue that the thing they are being paged about is not actually that bad. That argument is easier to avoid than to win. Treating an uncertain incident as the higher severity and correcting it in the postmortem removes the debate from the moment when cognitive resources are scarcest.

Pair severity with postmortem triggers

Severity levels decide response intensity. They should also decide which incidents get a written postmortem. Google’s SRE book lists common postmortem triggers: user-visible downtime or degradation beyond a certain threshold, data loss of any kind, on-call engineer intervention such as a release rollback or traffic rerouting, a resolution time above some threshold, and a monitoring failure that implies manual incident discovery. The same chapter states that it is important to define postmortem criteria before an incident occurs so that everyone knows when a postmortem is necessary.

For a five-person team, the practical move is to write the postmortem trigger into the SEV table. If SEV-1 and SEV-2 always get a postmortem, and SEV-3 gets one when the on-call engineer intervened manually or when monitoring failed to catch it, then the team does not have to decide after the fact whether the incident “deserved” a writeup. The decision was made in advance.

This also gives you the correction mechanism for over-classification. If an incident was treated as SEV-2 under the assume-the-worst rule and the postmortem shows it was really SEV-3, that is a data point about your thresholds. It is not a failure. It is the system working as designed.

What the sources do not tell you

None of the primary sources cited here recommend a specific number of SEV levels for a five-person team, and none of them say a single-customer outage is always one level or another. That determination depends on your own metrics: how many customers or accounts you have, what counts as core functionality, and what your current redundancy posture is. The sources give you the shape of a metric-driven ladder and a tie-breaker rule. The numbers are yours.

What the sources do support is the underlying discipline. Google’s SRE book describes on-call engineers as expected to triage a page and work toward resolution, possibly involving other team members and escalating as needed, and notes that paging events take priority over almost every other task including project work. It also names clear escalation paths, well-defined incident-management procedures, and a blameless postmortem culture as the most important on-call resources. A written SEV table is one of those well-defined procedures. It is not bureaucracy; it is the thing that lets a small rotation respond consistently without a standing incident-command hierarchy.

The SRE workbook adds that formulating rules about how to communicate and coordinate before disaster strikes allows the team to concentrate on resolving an incident when it occurs, and lists declaring incidents early and often among the basic principles of incident response. A SEV table that is specific enough to apply in under a minute is what makes early declaration cheap.

A practical exercise

Before the next incident, write your SEV table with specific percentages and named functionality. Decide which levels page and which levels wait for business hours. Decide which levels trigger a postmortem. Then test the table against the last three incidents your team handled. For each one, ask: under these definitions, what level would it have been, and would the response have been different? If the answer is that the table would have produced a response the team could not actually sustain — waking three people on a Tuesday night for something that meets your SEV-2 threshold, for example — then the threshold is wrong, not the team.

Revise and repeat. The table is a living document, and the postmortem is where it gets corrected. For a related exercise in writing the recovery steps before you need them, see Write the Recovery Checklist Before You Need It.

FAQ

How many SEV levels should a five-person team have?

There is no sourced answer to this. Atlassian’s guidance is that you factor in team size, on-call schedules, traffic patterns, and incident frequency when setting levels. Three or four levels is a common starting shape, but the right number is the one your team can apply consistently under pressure.

Should a single-customer outage ever page?

It depends on your written threshold. If the affected functionality is core for a meaningful share of your customers, or if the failure suggests a wider problem, it may meet your major-incident threshold. If it is a partial loss for a small subset, a metric-driven definition likely places it at SEV-3 or SEV-4. The sources do not prescribe a specific mapping.

What if we disagree about the severity during an incident?

PagerDuty’s rule is to treat it as the higher severity and review during the postmortem. That removes the debate from the moment when it is most costly and puts the correction where it belongs.

Do we need a postmortem for every SEV-2?

Google’s SRE book lists common postmortem triggers including user-visible downtime beyond a threshold, data loss, on-call intervention, resolution time above a threshold, and monitoring failure. Writing your own triggers into the SEV table before an incident means the team does not have to decide afterward.

Is severity the same as priority?

No. Atlassian defines severity as a measurement of impact and priority as a measurement of urgency. They often align, but not always. Severity drives response intensity; priority drives sequencing.

Every lean team has a version of this conversation. A new engineer joins on-call, opens the primary database, and asks why the service runs PostgreSQL instead of the managed alternative the rest of the stack uses. The answer exists. It lives in the head of the person who made the call, in a Slack thread from two years ago, and in a half-finished design doc that was never merged. When that person leaves, the answer leaves with them.

A one-page decision record is the smallest artifact that prevents this. It is not a design document, not a runbook, and not a postmortem. It is a dated, immutable note that captures one choice, the constraints that made it reasonable, and the conditions under which it should be revisited. For a team of two to fifteen engineers who are their own on-call rotation, it is the cheapest form of institutional memory available.

What a decision record is, and what it is not

The concept is not new. The Architectural Decision Record community defines an architectural decision as “a justified design choice that addresses a functional or non-functional requirement that is architecturally significant,” and an ADR as a record that “captures a single AD and its rationale.” The same source notes that the practice extends beyond architecture to “any decision record.” Michael Nygard’s 2011 blog post popularized the format, and the community maintains a comparison of seven templates.

What matters for a small team is the constraint on size. A decision record is not a design doc. It does not enumerate every alternative in depth, does not contain implementation steps, and does not get updated as the system evolves. It is a snapshot of reasoning at a point in time. If you find yourself writing more than a page, you are writing a design doc, and it will not get read during an incident.

It is also not a runbook. A runbook tells you how to restore the database. A decision record tells you why the database is PostgreSQL, what you gave up by choosing it, and what would have to change for the choice to be wrong. The two documents serve different readers at different moments. The runbook is read at 3 a.m. during a restore. The decision record is read during onboarding, during a migration debate, and during the postmortem that asks whether the original choice contributed to the incident.

The fields that earn their place on one page

A useful decision record for a lean team has seven fields. Each one exists because its absence causes a specific failure mode.

Title and date. A short imperative title (“Use PostgreSQL 16 as primary datastore for billing service”) and the date the decision was made. The date matters because it anchors the decision to a version of the system and a set of constraints that may no longer hold. PostgreSQL 18.6 is the current documented release as of this writing; a decision made against PostgreSQL 12 should be read with that gap in mind.

Status. One of: proposed, accepted, superseded by [link], or deprecated. This is the field that keeps the record honest. A decision record that is never superseded becomes folklore. A decision record with a status field can be retired cleanly.

Context. The operational reality at the time. Team size, on-call rotation, budget ceiling, compliance obligations, existing skills, and the scale the service was expected to handle. This is the field most often skipped and most often needed. “We chose PostgreSQL because it was free” is not context. “We chose PostgreSQL because the team had two engineers with production PostgreSQL experience, the managed alternative cost $X per month at our projected write volume, and we needed row-level security for a compliance requirement that the alternative did not support at the time” is context.

Decision. One or two sentences stating what was chosen. No hedging.

Alternatives considered. A short list with the reason each was rejected. This is where the record earns its keep during a future migration debate. If the team later considers moving to a managed service, the record shows whether the original objection still applies. If the objection was cost at a specific volume, and volume has changed, the record tells you the decision is ripe for revisit.

Consequences. What the team accepted by making this choice. Operational burden, backup and recovery responsibility, upgrade cadence, the need for a WAL archiving strategy, the need for a tested restore procedure. This is the field that connects the decision record to the recovery checklist. If the decision to run PostgreSQL yourself means you own point-in-time recovery, the decision record should say so, and the recovery checklist should reflect it.

Revisit triggers. The conditions under which this decision should be reopened. Specific, observable conditions. “If write volume exceeds X per second,” “if the team drops below two engineers with PostgreSQL experience,” “if the managed alternative adds row-level security at a cost below Y.” Without this field, the decision record is a historical document. With it, the record is a live input to planning.

Why PostgreSQL specifically makes this exercise worthwhile

PostgreSQL is a reasonable default for many small teams, and that is exactly why the reasoning behind choosing it is easy to lose. The choice feels obvious in retrospect, so nobody writes it down. Then a new engineer arrives from a shop that ran MySQL or a managed document store, and the obviousness is no longer shared.

The operational surface of PostgreSQL is large enough that the decision has real consequences. The PostgreSQL documentation devotes chapters to backup and restore, high availability and replication, reliability and the write-ahead log, and monitoring database activity. Each of those chapters represents work that a lean team either does or explicitly decides not to do. A decision record that says “we chose PostgreSQL” without saying which of those responsibilities the team accepted is incomplete.

Configuration is another area where the original reasoning matters. Tools like postgresqlco.nf exist because PostgreSQL’s configuration parameters are numerous and their interactions are non-obvious. A decision record does not need to capture every setting, but it should capture the constraint that drove the initial configuration posture. If the team chose to run with default settings because the workload was small and predictable, that is a decision. If the team tuned for a specific write pattern, that is a decision. Either way, the next person needs to know which one it was.

Where the record lives, and how it is found

A decision record that cannot be found during an incident is not a decision record. It is a file.

The storage location should be the same place engineers already look for project context. For most teams, that is the repository. A docs/decisions/ directory with files named YYYY-MM-DD-short-title.md sorts chronologically and is discoverable with a directory listing. The GitHub documentation on READMEs notes that a README is often the first item a visitor sees and that GitHub surfaces a README from the .github, root, or docs directory. A short pointer in the repository README to the decisions directory costs one line and saves a search.

Naming matters more than format. A file named database-choice.md is ambiguous after three database-related decisions. A file named 2024-03-11-postgres-primary-datastore.md is self-describing. The date prefix also makes supersession visible: a newer file with a related title is easy to spot.

Linking matters as much as naming. The decision record should link to the runbook that covers backup and restore, and the runbook should link back to the decision record. The postmortem template should include a field for “relevant decision records.” The onboarding checklist should include a pass through the decisions directory. Each of these links is cheap to add and expensive to reconstruct later.

How decision records interact with postmortems and runbooks

The Google SRE book chapter on postmortem culture describes a postmortem as “a written record of an incident, its impact, the actions taken to mitigate or resolve it, the root cause(s), and the follow-up actions to prevent the incident from recurring.” It also notes that postmortems are expected after any significant undesirable event, and that common triggers include user-visible downtime, data loss, on-call intervention, and monitoring failure.

A decision record is not a postmortem, but the two documents should reference each other. When an incident traces back to a database configuration choice, the postmortem’s root cause analysis should cite the decision record that established the configuration. When the postmortem produces an action item to change the configuration, that action item should produce a new decision record or a supersession of the old one.

The SRE book also emphasizes that blameless postmortems focus on “identifying the contributing causes of the incident without indicting any individual or team.” A decision record supports this by making the original reasoning visible. If a decision that seemed sound at the time contributed to an incident, the record shows what information was available when the decision was made. That is the difference between “who chose this” and “what did we know when we chose this.”

Runbooks and decision records have a simpler relationship. The runbook is procedural. The decision record is contextual. A runbook that says “run pgBackRest restore” is more useful when the reader understands why pgBackRest was chosen over the built-in pg_basebackup and WAL archiving. The decision record provides that context without cluttering the runbook.

A minimal review process

The failure mode of decision records in small teams is not that they are written badly. It is that they are written once and never revisited, so they drift out of alignment with the system. The fix is a review cadence that is small enough to survive a busy quarter.

One approach: review the decisions directory once per quarter, during an existing team meeting. For each record, ask one question: is the status still accurate? If yes, leave it. If no, either supersede it with a new record or mark it deprecated with a one-line note explaining why. This is a fifteen-minute exercise for a team with a dozen records.

A second approach: tie review to events rather than calendar. When a new engineer joins, they read the decisions directory as part of onboarding and flag any record that does not match what they observe. When a postmortem produces an action item that changes a previously decided approach, the postmortem owner writes the superseding record. When a revisit trigger fires, the person who notices writes the new record.

Both approaches work. The important thing is that supersession is a normal, low-ceremony act. A superseded record is not a mistake. It is a record of a decision that was correct under the constraints that existed when it was made.

What this looks like in practice

A team of six engineers runs a billing service on PostgreSQL 16, self-managed on EC2. The original decision was made eighteen months ago by an engineer who has since left. The decision record, written at the time, says the team chose self-managed PostgreSQL over Amazon RDS because the projected write volume at the time would have pushed RDS into a cost tier the budget could not absorb, and because the team had two engineers with production PostgreSQL experience.

The consequences section notes that the team accepted responsibility for WAL archiving, point-in-time recovery, and minor version upgrades. The revisit triggers include “if the team drops below two engineers with PostgreSQL experience” and “if RDS pricing for the required instance class falls below the current EC2 cost plus 20 percent.”

Eighteen months later, one of the two experienced engineers has moved to a different team. The revisit trigger has fired. The current team reads the decision record, sees that the original cost objection may no longer apply, and opens a new decision record to evaluate the migration. The old record is marked superseded. The new record cites the old one. The reasoning is preserved across the transition.

None of this required a design doc. It required one page, written at the time of the decision, by the person who made it.

Frequently asked questions

How is a decision record different from a design doc? A design doc describes how a system will be built. A decision record describes why a specific choice was made and what was traded away. Design docs are often long and are written before implementation. Decision records are short and are written at the moment of decision. A design doc may contain several decision records, or none.

What if the decision was made years ago and nobody remembers the reasoning? Write the record now, with the date of the original decision if known, and note in the context section that the reasoning is reconstructed. A reconstructed record is better than no record. It also creates a natural moment to ask whether the decision still holds.

Should every decision get a record? No. The ADR community frames the threshold as decisions that address “architecturally significant” requirements, meaning requirements with a measurable effect on architecture and quality. For a lean team, a practical filter is: if a new engineer would need to know this to avoid making a mistake, write it down. Database choice, backup strategy, authentication approach, and deployment topology usually qualify. Library version bumps usually do not.

How do decision records relate to the recovery checklist? The decision record explains why the recovery approach was chosen. The recovery checklist is the procedure. If the decision record says the team accepted responsibility for point-in-time recovery, the recovery checklist should include a tested PITR procedure. The two documents should link to each other.

What happens when a decision record is wrong? Mark it superseded, write a new record that explains what changed, and link the two. The old record stays in place. Its value is in showing what was known at the time, not in being correct forever.

Does this work for decisions other than database choice? Yes. The format is general. The adr.github.io project notes that ADR usage “can be extended to design and other decisions.” The constraint is the same: one page, dated, with context, consequences, and revisit triggers.

When a sole on-call engineer leaves a 2–15 person team, the loss is rarely the code. It is the unwritten operational knowledge: which backup actually restores, which alert is noise, which break-glass path still works, and which service has a history that never made it into a runbook. Two weeks is enough to extract the highest-value pieces if the process is structured around verification rather than conversation.

What is actually at risk

In lean teams, operational knowledge concentrates in the person who has been paged the most. The artifacts most commonly lost are not documents but decision context: why a particular pgBackRest stanza uses a specific retention policy, why a restic repository is mounted read-only on one host, why a PagerDuty escalation rule bypasses the primary on-call after 10 minutes. These are not recoverable from configuration alone because the configuration records the what, not the why.

The second category is procedural: the exact sequence of commands used during the last restore, including the flags that were not in the runbook. PostgreSQL’s continuous archiving documentation is explicit that recovery requires a continuous sequence of archived WAL files extending back at least as far as the start time of the base backup, and that the archive command must return zero exit status only on success. A departing engineer may know that the archive_command on one cluster silently returns zero on a pre-existing file, or that a particular base backup label is the only one known to restore cleanly. That knowledge is not in the configuration file.

The third category is social and historical: which vendor support contract covers which service, which internal team owns a shared dependency, which incident in the past six months produced a workaround that is still load-bearing. Google’s SRE book describes postmortems as a written record of an incident, its impact, the actions taken to mitigate or resolve it, the root cause, and follow-up actions. If those postmortems exist, they are a starting point. If they do not, the departing engineer is the only remaining copy.

Prioritize extraction by blast radius, not by topic

Two weeks is roughly ten working days. A useful allocation is to spend the first three days on backup and restore, the next three on access and break-glass, two on alerting and escalation, and the final two on verification and handoff. This ordering is not arbitrary: backup and restore failures are the only category where the team may not discover the gap until an actual data-loss event, and by then the departing engineer is gone.

For backup and restore, the extraction target is not a description of the backup system. It is a tested restore. The restic documentation describes restore as a command that can target a specific snapshot, a specific path within a snapshot, or the latest snapshot filtered by host and path. The departing engineer should walk through the exact command used in the last drill, including the repository path, the password source, and the target directory. If the command uses --include or --exclude, those patterns should be captured verbatim. If the restore is done in-place, the documentation warns that an interrupted restore can leave files in a partially restored state, so the team should know whether the last drill used a temporary target or an in-place restore.

For PostgreSQL, the extraction should cover the pgBackRest stanza configuration, the archive_command or archive_library setting, and the retention policy. pgBackRest’s user guide notes that a differential backup depends on the previous full backup, and an incremental backup depends on all prior incremental backups back to the prior differential or full backup. The departing engineer should identify which backup sets are known to be restorable and which are assumed to be restorable but never tested. The difference matters during an incident.

For access and break-glass, the extraction target is the path that works when normal authentication fails. This includes the location of emergency credentials, the conditions under which they are used, and the person or process that rotates them afterward. AWS IAM documentation covers delegation and credential management, but the team-specific question is which role or user is the actual break-glass identity, and whether it has been tested since the last rotation. If the departing engineer is the only person who has used it, that is a finding, not a fact to record.

For alerting and escalation, the extraction target is the mapping between symptom and action. PagerDuty’s on-call playbook was not retrievable for this article, so the guidance here is limited to what can be observed from the team’s own configuration: which alerts page, which alerts only notify, which escalation rules exist, and which alerts have been acknowledged without action in the past 90 days. The departing engineer should identify any alert that is known to be a false positive and any alert that has never fired but is expected to. The second category is the dangerous one.

Use structured prompts, not open-ended interviews

Open-ended knowledge transfer sessions tend to produce narrative, not procedure. A more effective format is a set of prompts tied to specific artifacts. For each backup repository, ask: what is the exact restore command, what is the expected duration, what is the failure mode if the repository is unavailable, and what is the last known-good snapshot. For each database cluster, ask: what is the stanza name, what is the retention policy, what is the archive command, and what is the last tested restore point. For each alert, ask: what does this alert mean, what is the first diagnostic step, and what is the escalation path if the first step does not resolve it.

These prompts should be answered in writing during the session, not after. The departing engineer should type the commands or paste the configuration while the remaining engineer watches. This produces a draft runbook that can be verified immediately. Google’s postmortem guidance emphasizes collaboration and knowledge-sharing at every stage, with real-time collaboration enabling rapid collection of data and ideas. The same principle applies to offboarding: the document should be built during the conversation, not reconstructed afterward.

For decision context, a useful prompt is: what would you do differently if you were designing this today, and why. This surfaces the tradeoffs that are not visible in the configuration. A pgBackRest stanza that uses a 30-day retention policy may be the result of a storage constraint that no longer exists, or a compliance requirement that still does. The departing engineer may be the only person who knows which. The answer should be recorded as a decision record, not as a runbook step, because it informs future changes rather than immediate operations.

Verify before access is revoked

The most common failure in offboarding knowledge extraction is that the remaining team believes it has captured the knowledge because it has written it down. The only reliable verification is execution by someone other than the departing engineer. For backup and restore, this means a restore drill during the two-week window, performed by the remaining engineer, with the departing engineer observing but not intervening. The drill should use the documented commands and the documented repository. If the restore fails, the failure is the finding, and the remaining time should be spent correcting the documentation and retesting.

For access and break-glass, verification means the remaining engineer uses the break-glass path to perform a low-risk action, such as listing a resource or reading a configuration value. If the path requires a credential that only the departing engineer holds, that is a finding. If the path requires a rotation step that is not documented, that is also a finding. The goal is not to test the credential’s power but to test the team’s ability to use it without the departing engineer.

For alerting and escalation, verification means the remaining engineer acknowledges a test alert and follows the documented escalation path. If the escalation path depends on a phone number or a schedule that only the departing engineer maintains, that dependency should be removed or reassigned before the last day. The verification should be recorded with a timestamp and the name of the person who performed it, so that the next offboarding has a baseline.

For runbooks and decision records, verification means the remaining engineer follows the runbook for a non-critical task, such as rotating a log file or checking a backup’s integrity. If the runbook requires a step that is not documented, the runbook is incomplete. The departing engineer should be available to answer questions during this verification, but the answers should be written into the runbook, not just spoken.

Integrate with access hygiene

Knowledge extraction and access revocation are not separate processes. The two-week window is also the window in which credentials should be rotated, break-glass paths should be reviewed, and least-privilege cleanup should be performed. The order matters: extraction should happen before revocation, but revocation should not wait until the last day. A useful pattern is to revoke non-essential access at the midpoint of the two weeks, after the backup and restore extraction is complete, and to revoke break-glass access only after the verification drill is complete.

This sequencing reduces the risk that the departing engineer’s access is used as a substitute for documentation. If the remaining engineer knows that the departing engineer can still restore the database, the incentive to verify the runbook is lower. Revoking access at the midpoint changes the incentive. It also surfaces dependencies earlier, when there is still time to resolve them.

Credential rotation should be treated as a verification step, not just a security step. If the departing engineer’s credentials are rotated and the remaining team can still perform all critical operations, the extraction is likely complete. If rotation breaks a critical operation, the extraction has a gap. The gap should be documented and resolved before the last day.

A lightweight process for a small team

A 2–15 person team does not need a dedicated offboarding program. It needs a repeatable sequence that fits in two weeks and produces artifacts that are useful after the departing engineer is gone. The sequence can be as simple as: day 1–3, backup and restore extraction with a restore drill; day 4–6, access and break-glass extraction with a break-glass drill; day 7–8, alerting and escalation extraction with a test alert; day 9–10, runbook and decision-record verification with a non-critical task. Each day should produce a written artifact and a verification record.

The artifacts should be stored where the remaining team already looks for operational information, not in a separate offboarding folder. A runbook that lives in a wiki page is more likely to be used than a runbook that lives in a shared drive. A decision record that lives next to the configuration it describes is more likely to be found than a decision record that lives in a meeting notes archive. The goal is to reduce the distance between the question and the answer.

The process should be reviewed after each offboarding. If a verification drill failed, the failure should be recorded and the process should be adjusted. If a runbook was incomplete, the gap should be noted and the template should be improved. Over time, the process becomes a form of institutional memory, not just a checklist. The Recovery Checklist Before You Need It is a useful companion for the backup and restore portion of this work, because it frames the drill as a pre-incident exercise rather than a post-departure scramble.

What cannot be captured in two weeks

Some knowledge resists extraction. Intuition about which alerts are noise, familiarity with the history of a service, and the social capital that comes from having worked with a vendor for years are not easily transferred. The goal of the two-week window is not to capture everything. It is to capture the knowledge that is both critical and verifiable: the restore command that works, the break-glass path that is tested, the escalation rule that is documented, and the decision record that explains why the configuration is the way it is.

The remaining knowledge should be treated as a known gap, not as a failure. A team that documents its gaps is more resilient than a team that assumes it has no gaps. The departing engineer should be asked to list the things they would want to know if they were staying, and that list should be added to the runbook as open questions. Some of those questions will be answered by the next incident. Some will be answered by the next hire. The important thing is that they are written down.

Frequently asked questions

How do we decide what to extract first when everything feels critical?

Start with the operations that have the longest recovery time and the least recent verification. A restore that has never been tested is higher priority than an alert that has been acknowledged without action for months. The blast radius of a failed restore is usually larger than the blast radius of a noisy alert.

What if the departing engineer is not available for the full two weeks?

Compress the schedule around the highest-blast-radius items. Backup and restore extraction can be done in a single day if the restore drill is scoped to a single repository and a single snapshot. Access and break-glass extraction can be done in half a day if the break-glass path is already documented. The verification steps are the ones that should not be skipped, because they are the only way to confirm that the extraction is accurate.

Should we record the extraction sessions?

Recording can be useful for reference, but it is not a substitute for written artifacts. A recording is difficult to search and easy to misinterpret. The written runbook and decision record should be the primary artifacts, with the recording as a backup. If the team uses recordings, they should be stored with the same access controls as the runbooks, because they may contain credential paths or internal hostnames.

How do we handle knowledge that is sensitive, such as break-glass credentials?

Sensitive knowledge should be documented in a way that describes the path without exposing the credential. The runbook should say where the credential is stored and how it is rotated, not what the credential is. The verification drill should confirm that the remaining engineer can access the credential through the documented path, not that the credential is written in the runbook.

What if the departing engineer is the only person who has ever performed a restore?

That is the highest-priority finding. The two-week window should be restructured to make the restore drill the first task, with the departing engineer observing and the remaining engineer executing. If the restore fails, the remaining time should be spent fixing the backup system, not documenting it. A documented restore that does not work is worse than an undocumented restore that does, because it creates false confidence.

How do we know when the extraction is complete?

The extraction is complete when the remaining team can perform all critical operations without the departing engineer’s access, and when the verification records show that those operations have been performed successfully. The absence of a departing engineer is not the test. The presence of a working runbook and a successful drill is the test.

Every new hire asks questions that the team has answered before. In a two-to-fifteen-engineer shop with no dedicated SRE function, those questions are the cheapest documentation audit you will ever run. The problem is that the answers usually live in a Slack thread, a pairing session, or someone’s memory, and the next hire asks the same thing six months later. A question log is a lightweight artifact that captures those questions, routes the ones that reveal real gaps into documentation, and leaves the rest alone.

This is not a knowledge-base project. It is a triage habit. The goal is to find the gaps before the next hire, not to document everything a new engineer might ever wonder about.

Why new-hire questions are a better audit than a documentation sprint

A documentation sprint asks the team to imagine what is missing. A new hire asks what is actually missing, in the order they encounter it. That order matters: it reflects the real path from laptop to production, not the org chart’s idea of how work flows.

The AWS Well-Architected Framework’s operational excellence pillar frames operations as something that should not be isolated from the teams it supports, and it recommends learning from operational events and sharing that learning. New-hire questions are an operational event with a predictable cadence. Treating them as a documentation input is consistent with that guidance, and it costs less than a dedicated review.

There is a second reason. In small teams, the people who answer questions are the same people who are on call. If the answer to “how do I rotate the age key for restic?” lives only in one engineer’s head, that engineer is a single point of failure for both documentation and incident response. Capturing the question and its answer is a small step toward reducing that dependency.

What belongs in the log

A question log is a list of questions, not a list of answers. The answer can be a link, a command, or a short paragraph. The question is the durable part because it tells you what the next person will need.

Useful fields:

  • Date asked. Helps you see whether a gap keeps recurring.
  • Question, in the asker’s words. Do not rewrite it into a documentation title yet. The raw phrasing is a signal about how people search.
  • Where it was asked. Slack channel, pairing session, pull request comment, or incident channel. This tells you where the answer already lives.
  • Answer or pointer. A link to the runbook, a command, or a note that no answer exists.
  • Gap? A yes/no judgment made during triage, not by the asker.
  • Action. One of: document, link from an existing doc, add to onboarding checklist, or no action.

The log can live in a shared document, a pinned Slack thread, or an issue tracker. The format matters less than the triage. A spreadsheet with a weekly review beats a beautiful wiki page that nobody reads.

Separating real gaps from normal onboarding friction

Not every question is a documentation gap. Some questions are the normal cost of learning a system, and documenting them would produce noise. The distinction is whether the answer is stable, discoverable, and needed by more than one person.

A question is probably a real gap when:

  • The answer exists but is not written down anywhere searchable.
  • The answer is written down but in a place the new hire could not reasonably find.
  • The answer requires a decision that is not recorded, such as why a particular backup tool was chosen.
  • Two different people give different answers.
  • The question is about a procedure that has operational consequences, such as restoring a database or rotating credentials.

A question is probably not a gap when:

  • It is about a one-off task that will not recur.
  • It is answered by reading the code, and the code is the source of truth.
  • It is a preference question with no operational consequence.
  • The answer is already in the onboarding checklist and the new hire simply had not reached that step.

The judgment is the hard part. The person doing triage should be someone who knows the system well enough to tell the difference, and who is not the new hire. In a small team, that is usually the on-call lead or the engineer who owns the relevant runbook.

How to run the log without adding process overhead

The failure mode of a question log is that it becomes a second job. The way to avoid that is to keep capture cheap and triage bounded.

Capture. Ask the new hire to add questions to the log as they arise, not after the fact. A single shared document with a table is enough. If the team already uses an issue tracker, a label like onboarding-question works. The important thing is that the new hire does not have to decide whether a question is worth logging. Log everything; triage later.

Triage. Once a week, or at the end of the new hire’s first month, spend fifteen minutes reviewing the log. For each question, mark it as a gap or not, and assign one action. The action should be small: add a paragraph to an existing runbook, add a link from the onboarding checklist, or open a ticket to write a new document. Do not try to write the document during triage.

Act. The action items go into the normal work queue. In a small team, that means they compete with feature work and incident follow-up. That is fine. The point is that they are visible and can be prioritized, not that they are done immediately.

Close the loop. When a gap is documented, add a link to the answer in the log. When the next hire asks the same question, the answer should be findable. If it is not, the documentation did not solve the problem.

What to document first

Not all gaps are equal. In a lean team running production without dedicated reliability coverage, the highest-value documentation is the kind that reduces the blast radius of an incident or the time to recover.

Prioritize questions about:

  • Backup and restore. How to restore a restic snapshot, how to verify a pgBackRest backup, how to find the right WAL segment for point-in-time recovery. The PostgreSQL documentation on continuous archiving and point-in-time recovery is explicit that you need a continuous sequence of archived WAL files extending back at least as far as the start time of your base backup, and that you should set up and test your archiving procedure before taking your first base backup. A new hire asking “how do I know the WAL archive is working?” is pointing at a real operational gap.
  • Access and break-glass. How to get production access, how to use the break-glass path, and what to do afterward. These questions often reveal that the process is undocumented or that the documentation is stale.
  • Incident response. Who to page, how to escalate, where the incident channel is, and what the postmortem process looks like. PagerDuty’s public incident response documentation is used to prepare new employees for on-call responsibilities, which is a useful model for what a new hire needs to know before their first shift.
  • Monitoring and alerts. What the alerts mean, which ones are actionable, and which ones are noise. A new hire asking “why did this alert fire?” may be pointing at an alert that needs tuning.

Questions about local development setup, code style, and tool preferences are lower priority unless they recur across hires.

Using the log to improve onboarding, not just documentation

The log has a second use: it tells you what the onboarding checklist is missing. If three consecutive hires ask how to request database access, the checklist is wrong, not the documentation.

Review the log at the end of each onboarding cycle and update the checklist. This is a small change that compounds. It also gives the new hire a sense that their questions are improving the team’s process, which is consistent with the blameless, learning-oriented culture described in Google’s SRE book. The book notes that postmortems are expected after significant undesirable events, and that the goal is to understand contributing causes without indicting individuals. The same posture applies to onboarding questions: the question is not a failure by the new hire, it is a signal about the system.

Metrics that tell you the log is working

You do not need a dashboard. A few simple counts are enough:

  • Repeat questions. How many questions in this onboarding cycle were already in the log from a previous cycle? If the number is going down, documentation is working. If it is flat or going up, the answers are not discoverable.
  • Time to first independent change. How long before the new hire shipped a change without asking for help? This is a rough measure, but it is the outcome the log is meant to improve.
  • Gap closure rate. How many gaps marked during triage were actually documented? A low rate means the actions are not being prioritized.
  • Questions about incidents. How many questions are about how to respond to a specific failure mode? These are the highest-value gaps because they affect recovery time.

None of these are precise. They are directional. The point is to notice when the log is filling up with the same questions and nothing is changing.

A worked example

Suppose a new hire joins a team running PostgreSQL on a single primary with pgBackRest and restic for file-level backups. In their first two weeks, they ask:

  1. How do I restore a single table from a pgBackRest backup?
  2. Where are the restic repository credentials stored?
  3. How do I know the WAL archive is current?
  4. What is the difference between the staging and production alert channels?
  5. How do I request access to the production database?

During triage, the team marks questions 1, 3, and 5 as gaps. Question 2 is answered by a link to the secrets manager documentation, which the new hire had not found. Question 4 is answered by a link to the monitoring runbook, which already existed but was not linked from the onboarding checklist.

The actions are: add a restore example to the pgBackRest runbook, add a WAL archive check to the daily operations checklist, and add the access request process to the onboarding checklist. None of these are large. Together they close three gaps that would otherwise have been rediscovered by the next hire.

What not to do

Do not turn the log into a documentation backlog that grows faster than it shrinks. If a question is marked as a gap and not documented within a reasonable time, either the gap is not important or the team does not have capacity. Both are useful signals. A backlog of fifty undocumented gaps is worse than a log of five that get closed.

Do not ask the new hire to write the documentation. They do not yet know enough to write it accurately, and the exercise can feel like busywork. The new hire’s job is to ask questions and log them. The team’s job is to answer and document.

Do not document everything. The goal is to reduce repeat questions and improve incident response, not to produce a complete manual. A lean team’s documentation should be a map of the terrain that matters, not a transcript of every conversation.

Where this fits with other resilience practices

A question log is not a substitute for a recovery drill or a postmortem. It is a complement. The recovery checklist you write before you need it, the postmortem you write after an incident, and the question log you keep during onboarding all serve the same purpose: they move knowledge from one person’s head into a form the team can use.

If you are starting from nothing, start with the recovery checklist. It is the highest-stakes document a lean team can have. Then add the question log. It is the cheapest way to find the next gap.

FAQ

How long should we keep the question log?

Keep it as long as it is useful. A log that spans several onboarding cycles is more valuable than one that is reset each time, because it shows which questions recur. If the log becomes too large to review in fifteen minutes, archive the closed items and keep the open gaps.

Who should own triage?

The person who owns the relevant runbook or the on-call lead. In a small team, that is often the same person. The owner should not be the new hire, because the judgment about what is a real gap requires system knowledge.

What if the team does not have time to document the gaps?

Then the log is telling you something about capacity. Either the gaps are not important enough to prioritize, or the team is too stretched to improve its own documentation. Both are worth knowing. The log does not create the problem; it makes it visible.

Should the log be public to the whole team?

Yes, unless it contains sensitive information. Questions about access paths and break-glass procedures may need to be restricted. In that case, keep the question in the log and put the answer in the restricted runbook.

How is this different from a runbook?

A runbook is a procedure. A question log is a list of things people did not know. The log feeds the runbook. It is the input, not the output.

Offboarding is a reliability event disguised as an HR event. When an engineer leaves a 2–15 person team, the same person who held production access, carried a pager, and knew which runbook to open at 03:00 is now a set of credentials that must be revoked. The failure mode is not that revocation is forgotten; it is that revocation is done in the wrong order, and the on-call rotation loses a responder, a break-glass path, or a backup credential at the moment it is needed.

This checklist is written for teams that are their own SRE function. It assumes AWS, GCP, or bare metal, PostgreSQL with WAL archiving, restic or pgBackRest for backups, and PagerDuty or an equivalent escalation tool. It is a sequence, not a policy document. The goal is to remove a person’s access within 48 hours while keeping the rotation intact.

Why 48 hours, and why order matters

NIST SP 800-53 Rev. 5, the control catalog for federal information systems, includes an Access Control family covering account management. The operational reading for a lean team is that revocation must be a defined procedure with an owner, not an ad hoc action. The 48-hour window is a team-level target, not a NIST requirement. Treat it as a service-level objective for your own process.

Order matters because access is layered. A departing engineer may hold a long-lived IAM user with an access key, a federated identity through an IdP, a PagerDuty user with a schedule override, a database role with a password in a password manager, and an SSH key on a bastion. Revoking the IAM user first can break a script that the rotation depends on. Revoking the PagerDuty user first can leave a gap in the escalation policy. The sequence below removes human access first, then service access, then shared secrets, then the pager identity.

Phase 1: Freeze and inventory (0–4 hours)

Before removing anything, capture what exists. This is the step that prevents the 03:00 surprise.

  • List all IAM users, roles, and access keys associated with the person in AWS. In AWS, temporary credentials issued through AWS STS are short-lived and expire on their own; long-lived IAM user access keys do not. The AWS IAM documentation states that temporary credentials can be configured to last from a few minutes to several hours and that after they expire, AWS no longer recognizes them. If the departing engineer’s access is federated through an IdP, the long-lived key inventory may be smaller than expected.
  • List all GCP principals. Google Cloud IAM grants roles to principals on resources, and allow policies are inherited down the resource hierarchy. A role granted at the organization or folder level applies to every descendant resource. Revoking a project-level binding does not remove an inherited organization-level binding. The Google Cloud IAM overview documents this inheritance explicitly: to understand who can access a resource, you must view the resource’s allow policy and its ancestors’ allow policies.
  • List PagerDuty schedule memberships, escalation policy steps, and any user-level notification rules. The PagerDuty integrations documentation describes the platform as connecting tools and data to incident management; the operational point is that a user is a first-class object in schedules and escalation policies, not just an email address.
  • List database roles and authentication methods. For PostgreSQL, check pg_hba.conf for any trust or md5 entries tied to a specific host or user, and check for roles with SUPERUSER or REPLICATION attributes.
  • List backup credentials. If restic repositories are accessed with a password stored in a shared vault, note who knows it. If pgBackRest uses an SSH key or a cloud identity, note which one.

Write the inventory to a ticket or a decision record. The inventory is the artifact that makes the rest of the checklist auditable.

Phase 2: Remove human access (4–24 hours)

Human access is the interactive path: console logins, CLI sessions, SSH sessions, and database sessions. Remove it before touching service accounts.

AWS

If the engineer has an IAM user, deactivate the access keys and delete the login profile. If the engineer accesses AWS through a federated identity, disable the identity in the IdP. The AWS documentation notes that temporary credentials are the basis for roles and identity federation and that they do not have to be explicitly revoked when no longer needed because they expire. That property is useful during offboarding: a federated session that is already issued will expire on its own, but the IdP must be disabled so new sessions cannot be issued.

Check for any IAM roles with a trust policy that allows the departing engineer’s principal to assume them. A role trust policy is not a permission grant to the person, but it is an access path. Remove the principal from the trust policy if the role was created for that person.

GCP

Remove the principal from all allow policies where it appears. Because of policy inheritance, check the organization, folder, and project levels. The Google Cloud IAM overview notes that the IAM API is eventually consistent: a write followed immediately by a read may return an older version, and changes may take time to affect access checks. Do not assume a revocation is effective the moment the API returns success. Verify with a read after a short delay, or use the policy troubleshooter.

If the engineer has a service account key, delete the key. Service account keys are long-lived credentials and are a common offboarding gap because they are often created for automation and then used interactively.

Bare metal and SSH

Remove the public key from authorized_keys on every host the engineer could reach. If you use a configuration management tool, remove the key from the source of truth and let the tool converge. If you use a bastion with short-lived certificates, revoke the certificate authority’s trust for that principal if the CA supports it; otherwise, wait for the certificate to expire and confirm the expiry window.

PostgreSQL

Revoke login on the role: ALTER ROLE username NOLOGIN;. Do not drop the role immediately if it owns objects or has granted privileges; dropping a role that owns objects requires reassigning ownership first. Check pg_shdepend for dependencies. If the role has SUPERUSER, revoke it before disabling login so the change is visible in the audit log.

If the engineer had a .pgpass file or a password in a shared vault, rotate the password for any role that was shared. A shared role is a single point of failure during offboarding; if two engineers used the same database role, disabling one person’s login does not remove the other person’s access, but it also does not remove the departing engineer’s access if the password is still known.

Phase 3: Service accounts and automation (24–36 hours)

Service accounts are where offboarding breaks the rotation. A departing engineer may have created a service account for a synthetic check, a backup job, or a deployment pipeline. If that service account is disabled without replacement, the check stops running and the rotation loses a signal.

For each service account the engineer owned or used:

  • Identify what depends on it. Check CI/CD pipelines, cron jobs, Kubernetes service accounts, and monitoring integrations.
  • Create a replacement service account with the minimum permissions needed, or transfer ownership to a shared team identity.
  • Update the dependency to use the replacement.
  • Disable the old service account and delete its keys.

In GCP, service accounts are principals. The IAM overview distinguishes human users from workloads and notes that service accounts are workload principals. That distinction is useful: a service account should not be used interactively, and an interactive user should not be the sole owner of a service account.

In AWS, the equivalent is an IAM role for EC2 or a role assumed by a pipeline. The AWS documentation describes roles for Amazon EC2 as a way to provide temporary credentials to instances without storing long-term credentials. If the departing engineer created an IAM user with an access key for a script, replace it with a role.

Phase 4: Backups and recovery paths (36–44 hours)

Backup access is the most commonly missed offboarding item because it is often shared. The question is not whether the departing engineer can restore; it is whether the remaining team can restore without them.

PostgreSQL WAL archiving

PostgreSQL’s continuous archiving documentation describes WAL archiving as a sequence of segment files copied by an archive_command or an archive_library. The archive command runs under the ownership of the PostgreSQL server user. If the archive command depends on a credential that the departing engineer controlled, the archive will fail after offboarding. Check the command for references to a specific home directory, an SSH key, or a cloud credential file.

The documentation also notes that if the archive command fails repeatedly, the pg_wal/ directory will fill with segment files, and if the file system fills, PostgreSQL will perform a PANIC shutdown. That is a production outage caused by an offboarding gap. Verify that the archive command uses a credential owned by the service, not by a person.

Test the restore path. A backup that has never been restored is a hypothesis. The Gray Haven article Write the Recovery Checklist Before You Need It describes the practice of writing the recovery procedure before an incident; use it here to confirm that the remaining team can perform a point-in-time recovery without the departing engineer’s credentials.

restic and pgBackRest

If restic repositories are encrypted with a password that the departing engineer knows, rotate the repository password. restic supports multiple keys per repository; add a new key for the team, then remove the old key. If pgBackRest uses a repository host with SSH keys, replace the key and update the configuration.

If the backup repository is in cloud storage, check the bucket or container policy for any principal that matches the departing engineer. In GCP, a bucket policy is an allow policy on the bucket resource; inherited policies from the project may also apply. In AWS, check the bucket policy and any IAM role that grants access.

Phase 5: PagerDuty and the on-call rotation (44–48 hours)

The pager identity is the last thing to remove because it is the thing that keeps the rotation working. Removing a user from PagerDuty before updating the schedule can create a gap; updating the schedule before removing the user can leave a stale notification path.

Sequence:

  1. Add a replacement responder to every schedule the departing engineer was on. If the team is small, this may mean the remaining engineers absorb the shift. Confirm that the replacement has the same escalation policy coverage.
  2. Update escalation policies to remove the departing engineer as a step. If the engineer was the only person in a step, replace the step before removing the user.
  3. Check for schedule overrides. A departing engineer may have an override that extends beyond their last day.
  4. Remove the user from PagerDuty. If the user is also the account owner or billing contact, transfer those roles first.
  5. Verify that a test incident routes to the remaining responders. PagerDuty’s integrations documentation describes email and API integrations; use a test event through the same path your production alerts use.

If the departing engineer was the only person who knew how to acknowledge a specific alert, that is a runbook gap, not a PagerDuty gap. Fix the runbook before the last day.

Break-glass access

Break-glass access is the credential you use when normal access is unavailable. It is also the credential most likely to be tied to a person. If the departing engineer held the break-glass credential, rotate it. If the break-glass credential is a shared password in a vault, change the password and update the vault entry.

The CISA Zero Trust Maturity Model page was not retrievable at the time of writing, so this article does not cite it. The operational principle is narrower: break-glass access should be auditable and time-bound. If your break-glass path is a long-lived IAM user with an access key, replace it with a role that requires MFA and issues temporary credentials. AWS documents temporary credentials as short-term and expiring; that property limits the window in which a lost break-glass credential can be used.

What to verify before closing the ticket

  • The departing engineer cannot authenticate to AWS, GCP, or any bare-metal host.
  • No service account or automation job depends on a credential the engineer controlled.
  • WAL archiving is still running and the archive command does not reference a personal credential.
  • A restore drill has been performed by a remaining team member within the last quarter.
  • The PagerDuty schedule has no gaps and a test incident routes correctly.
  • Break-glass credentials have been rotated.
  • The inventory and the actions taken are recorded in a decision record or ticket.

FAQ

Do we need to revoke temporary credentials explicitly?

No. AWS documents temporary credentials as expiring after a configured lifetime, after which AWS no longer recognizes them. The action required during offboarding is to prevent new temporary credentials from being issued, which means disabling the identity in the IdP or removing the principal from the role trust policy.

What if the departing engineer is the only person who knows the backup password?

That is a single point of failure that predates the offboarding. Rotate the password before the last day, add a new key for the team, and verify a restore. If the password is lost, the backup may be unrecoverable; treat that as an incident and document it.

How do we handle a shared database role?

Replace it with individual roles. A shared role cannot be revoked for one person without affecting others. If individual roles are not feasible, rotate the password and distribute it only to the remaining team, then plan to split the role.

What if the engineer leaves before we finish the checklist?

Prioritize human access removal and PagerDuty coverage. Service accounts and backup credentials can be rotated after the last day, but an active human credential and an uncovered pager are immediate risks. If the engineer is cooperative, ask them to transfer ownership of service accounts before their last day.

Does NIST SP 800-53 require a specific revocation timeframe?

No. SP 800-53 Rev. 5 includes an Access Control family covering account management, but the specific requirements and any timeframe are organizational decisions. The 48-hour window in this article is a team-level target, not a NIST requirement.

Sources