When a sole on-call engineer leaves a 2–15 person team, the loss is rarely the code. It is the unwritten operational knowledge: which backup actually restores, which alert is noise, which break-glass path still works, and which service has a history that never made it into a runbook. Two weeks is enough to extract the highest-value pieces if the process is structured around verification rather than conversation.
What is actually at risk
In lean teams, operational knowledge concentrates in the person who has been paged the most. The artifacts most commonly lost are not documents but decision context: why a particular pgBackRest stanza uses a specific retention policy, why a restic repository is mounted read-only on one host, why a PagerDuty escalation rule bypasses the primary on-call after 10 minutes. These are not recoverable from configuration alone because the configuration records the what, not the why.
The second category is procedural: the exact sequence of commands used during the last restore, including the flags that were not in the runbook. PostgreSQL’s continuous archiving documentation is explicit that recovery requires a continuous sequence of archived WAL files extending back at least as far as the start time of the base backup, and that the archive command must return zero exit status only on success. A departing engineer may know that the archive_command on one cluster silently returns zero on a pre-existing file, or that a particular base backup label is the only one known to restore cleanly. That knowledge is not in the configuration file.
The third category is social and historical: which vendor support contract covers which service, which internal team owns a shared dependency, which incident in the past six months produced a workaround that is still load-bearing. Google’s SRE book describes postmortems as a written record of an incident, its impact, the actions taken to mitigate or resolve it, the root cause, and follow-up actions. If those postmortems exist, they are a starting point. If they do not, the departing engineer is the only remaining copy.
Prioritize extraction by blast radius, not by topic
Two weeks is roughly ten working days. A useful allocation is to spend the first three days on backup and restore, the next three on access and break-glass, two on alerting and escalation, and the final two on verification and handoff. This ordering is not arbitrary: backup and restore failures are the only category where the team may not discover the gap until an actual data-loss event, and by then the departing engineer is gone.
For backup and restore, the extraction target is not a description of the backup system. It is a tested restore. The restic documentation describes restore as a command that can target a specific snapshot, a specific path within a snapshot, or the latest snapshot filtered by host and path. The departing engineer should walk through the exact command used in the last drill, including the repository path, the password source, and the target directory. If the command uses --include or --exclude, those patterns should be captured verbatim. If the restore is done in-place, the documentation warns that an interrupted restore can leave files in a partially restored state, so the team should know whether the last drill used a temporary target or an in-place restore.
For PostgreSQL, the extraction should cover the pgBackRest stanza configuration, the archive_command or archive_library setting, and the retention policy. pgBackRest’s user guide notes that a differential backup depends on the previous full backup, and an incremental backup depends on all prior incremental backups back to the prior differential or full backup. The departing engineer should identify which backup sets are known to be restorable and which are assumed to be restorable but never tested. The difference matters during an incident.
For access and break-glass, the extraction target is the path that works when normal authentication fails. This includes the location of emergency credentials, the conditions under which they are used, and the person or process that rotates them afterward. AWS IAM documentation covers delegation and credential management, but the team-specific question is which role or user is the actual break-glass identity, and whether it has been tested since the last rotation. If the departing engineer is the only person who has used it, that is a finding, not a fact to record.
For alerting and escalation, the extraction target is the mapping between symptom and action. PagerDuty’s on-call playbook was not retrievable for this article, so the guidance here is limited to what can be observed from the team’s own configuration: which alerts page, which alerts only notify, which escalation rules exist, and which alerts have been acknowledged without action in the past 90 days. The departing engineer should identify any alert that is known to be a false positive and any alert that has never fired but is expected to. The second category is the dangerous one.
Use structured prompts, not open-ended interviews
Open-ended knowledge transfer sessions tend to produce narrative, not procedure. A more effective format is a set of prompts tied to specific artifacts. For each backup repository, ask: what is the exact restore command, what is the expected duration, what is the failure mode if the repository is unavailable, and what is the last known-good snapshot. For each database cluster, ask: what is the stanza name, what is the retention policy, what is the archive command, and what is the last tested restore point. For each alert, ask: what does this alert mean, what is the first diagnostic step, and what is the escalation path if the first step does not resolve it.
These prompts should be answered in writing during the session, not after. The departing engineer should type the commands or paste the configuration while the remaining engineer watches. This produces a draft runbook that can be verified immediately. Google’s postmortem guidance emphasizes collaboration and knowledge-sharing at every stage, with real-time collaboration enabling rapid collection of data and ideas. The same principle applies to offboarding: the document should be built during the conversation, not reconstructed afterward.
For decision context, a useful prompt is: what would you do differently if you were designing this today, and why. This surfaces the tradeoffs that are not visible in the configuration. A pgBackRest stanza that uses a 30-day retention policy may be the result of a storage constraint that no longer exists, or a compliance requirement that still does. The departing engineer may be the only person who knows which. The answer should be recorded as a decision record, not as a runbook step, because it informs future changes rather than immediate operations.
Verify before access is revoked
The most common failure in offboarding knowledge extraction is that the remaining team believes it has captured the knowledge because it has written it down. The only reliable verification is execution by someone other than the departing engineer. For backup and restore, this means a restore drill during the two-week window, performed by the remaining engineer, with the departing engineer observing but not intervening. The drill should use the documented commands and the documented repository. If the restore fails, the failure is the finding, and the remaining time should be spent correcting the documentation and retesting.
For access and break-glass, verification means the remaining engineer uses the break-glass path to perform a low-risk action, such as listing a resource or reading a configuration value. If the path requires a credential that only the departing engineer holds, that is a finding. If the path requires a rotation step that is not documented, that is also a finding. The goal is not to test the credential’s power but to test the team’s ability to use it without the departing engineer.
For alerting and escalation, verification means the remaining engineer acknowledges a test alert and follows the documented escalation path. If the escalation path depends on a phone number or a schedule that only the departing engineer maintains, that dependency should be removed or reassigned before the last day. The verification should be recorded with a timestamp and the name of the person who performed it, so that the next offboarding has a baseline.
For runbooks and decision records, verification means the remaining engineer follows the runbook for a non-critical task, such as rotating a log file or checking a backup’s integrity. If the runbook requires a step that is not documented, the runbook is incomplete. The departing engineer should be available to answer questions during this verification, but the answers should be written into the runbook, not just spoken.
Integrate with access hygiene
Knowledge extraction and access revocation are not separate processes. The two-week window is also the window in which credentials should be rotated, break-glass paths should be reviewed, and least-privilege cleanup should be performed. The order matters: extraction should happen before revocation, but revocation should not wait until the last day. A useful pattern is to revoke non-essential access at the midpoint of the two weeks, after the backup and restore extraction is complete, and to revoke break-glass access only after the verification drill is complete.
This sequencing reduces the risk that the departing engineer’s access is used as a substitute for documentation. If the remaining engineer knows that the departing engineer can still restore the database, the incentive to verify the runbook is lower. Revoking access at the midpoint changes the incentive. It also surfaces dependencies earlier, when there is still time to resolve them.
Credential rotation should be treated as a verification step, not just a security step. If the departing engineer’s credentials are rotated and the remaining team can still perform all critical operations, the extraction is likely complete. If rotation breaks a critical operation, the extraction has a gap. The gap should be documented and resolved before the last day.
A lightweight process for a small team
A 2–15 person team does not need a dedicated offboarding program. It needs a repeatable sequence that fits in two weeks and produces artifacts that are useful after the departing engineer is gone. The sequence can be as simple as: day 1–3, backup and restore extraction with a restore drill; day 4–6, access and break-glass extraction with a break-glass drill; day 7–8, alerting and escalation extraction with a test alert; day 9–10, runbook and decision-record verification with a non-critical task. Each day should produce a written artifact and a verification record.
The artifacts should be stored where the remaining team already looks for operational information, not in a separate offboarding folder. A runbook that lives in a wiki page is more likely to be used than a runbook that lives in a shared drive. A decision record that lives next to the configuration it describes is more likely to be found than a decision record that lives in a meeting notes archive. The goal is to reduce the distance between the question and the answer.
The process should be reviewed after each offboarding. If a verification drill failed, the failure should be recorded and the process should be adjusted. If a runbook was incomplete, the gap should be noted and the template should be improved. Over time, the process becomes a form of institutional memory, not just a checklist. The Recovery Checklist Before You Need It is a useful companion for the backup and restore portion of this work, because it frames the drill as a pre-incident exercise rather than a post-departure scramble.
What cannot be captured in two weeks
Some knowledge resists extraction. Intuition about which alerts are noise, familiarity with the history of a service, and the social capital that comes from having worked with a vendor for years are not easily transferred. The goal of the two-week window is not to capture everything. It is to capture the knowledge that is both critical and verifiable: the restore command that works, the break-glass path that is tested, the escalation rule that is documented, and the decision record that explains why the configuration is the way it is.
The remaining knowledge should be treated as a known gap, not as a failure. A team that documents its gaps is more resilient than a team that assumes it has no gaps. The departing engineer should be asked to list the things they would want to know if they were staying, and that list should be added to the runbook as open questions. Some of those questions will be answered by the next incident. Some will be answered by the next hire. The important thing is that they are written down.
Frequently asked questions
How do we decide what to extract first when everything feels critical?
Start with the operations that have the longest recovery time and the least recent verification. A restore that has never been tested is higher priority than an alert that has been acknowledged without action for months. The blast radius of a failed restore is usually larger than the blast radius of a noisy alert.
What if the departing engineer is not available for the full two weeks?
Compress the schedule around the highest-blast-radius items. Backup and restore extraction can be done in a single day if the restore drill is scoped to a single repository and a single snapshot. Access and break-glass extraction can be done in half a day if the break-glass path is already documented. The verification steps are the ones that should not be skipped, because they are the only way to confirm that the extraction is accurate.
Should we record the extraction sessions?
Recording can be useful for reference, but it is not a substitute for written artifacts. A recording is difficult to search and easy to misinterpret. The written runbook and decision record should be the primary artifacts, with the recording as a backup. If the team uses recordings, they should be stored with the same access controls as the runbooks, because they may contain credential paths or internal hostnames.
How do we handle knowledge that is sensitive, such as break-glass credentials?
Sensitive knowledge should be documented in a way that describes the path without exposing the credential. The runbook should say where the credential is stored and how it is rotated, not what the credential is. The verification drill should confirm that the remaining engineer can access the credential through the documented path, not that the credential is written in the runbook.
What if the departing engineer is the only person who has ever performed a restore?
That is the highest-priority finding. The two-week window should be restructured to make the restore drill the first task, with the departing engineer observing and the remaining engineer executing. If the restore fails, the remaining time should be spent fixing the backup system, not documenting it. A documented restore that does not work is worse than an undocumented restore that does, because it creates false confidence.
How do we know when the extraction is complete?
The extraction is complete when the remaining team can perform all critical operations without the departing engineer’s access, and when the verification records show that those operations have been performed successfully. The absence of a departing engineer is not the test. The presence of a working runbook and a successful drill is the test.


