When a sole on-call engineer leaves a 2–15 person team, the loss is rarely the code. It is the unwritten operational knowledge: which backup actually restores, which alert is noise, which break-glass path still works, and which service has a history that never made it into a runbook. Two weeks is enough to extract the highest-value pieces if the process is structured around verification rather than conversation.

What is actually at risk

In lean teams, operational knowledge concentrates in the person who has been paged the most. The artifacts most commonly lost are not documents but decision context: why a particular pgBackRest stanza uses a specific retention policy, why a restic repository is mounted read-only on one host, why a PagerDuty escalation rule bypasses the primary on-call after 10 minutes. These are not recoverable from configuration alone because the configuration records the what, not the why.

The second category is procedural: the exact sequence of commands used during the last restore, including the flags that were not in the runbook. PostgreSQL’s continuous archiving documentation is explicit that recovery requires a continuous sequence of archived WAL files extending back at least as far as the start time of the base backup, and that the archive command must return zero exit status only on success. A departing engineer may know that the archive_command on one cluster silently returns zero on a pre-existing file, or that a particular base backup label is the only one known to restore cleanly. That knowledge is not in the configuration file.

The third category is social and historical: which vendor support contract covers which service, which internal team owns a shared dependency, which incident in the past six months produced a workaround that is still load-bearing. Google’s SRE book describes postmortems as a written record of an incident, its impact, the actions taken to mitigate or resolve it, the root cause, and follow-up actions. If those postmortems exist, they are a starting point. If they do not, the departing engineer is the only remaining copy.

Prioritize extraction by blast radius, not by topic

Two weeks is roughly ten working days. A useful allocation is to spend the first three days on backup and restore, the next three on access and break-glass, two on alerting and escalation, and the final two on verification and handoff. This ordering is not arbitrary: backup and restore failures are the only category where the team may not discover the gap until an actual data-loss event, and by then the departing engineer is gone.

For backup and restore, the extraction target is not a description of the backup system. It is a tested restore. The restic documentation describes restore as a command that can target a specific snapshot, a specific path within a snapshot, or the latest snapshot filtered by host and path. The departing engineer should walk through the exact command used in the last drill, including the repository path, the password source, and the target directory. If the command uses --include or --exclude, those patterns should be captured verbatim. If the restore is done in-place, the documentation warns that an interrupted restore can leave files in a partially restored state, so the team should know whether the last drill used a temporary target or an in-place restore.

For PostgreSQL, the extraction should cover the pgBackRest stanza configuration, the archive_command or archive_library setting, and the retention policy. pgBackRest’s user guide notes that a differential backup depends on the previous full backup, and an incremental backup depends on all prior incremental backups back to the prior differential or full backup. The departing engineer should identify which backup sets are known to be restorable and which are assumed to be restorable but never tested. The difference matters during an incident.

For access and break-glass, the extraction target is the path that works when normal authentication fails. This includes the location of emergency credentials, the conditions under which they are used, and the person or process that rotates them afterward. AWS IAM documentation covers delegation and credential management, but the team-specific question is which role or user is the actual break-glass identity, and whether it has been tested since the last rotation. If the departing engineer is the only person who has used it, that is a finding, not a fact to record.

For alerting and escalation, the extraction target is the mapping between symptom and action. PagerDuty’s on-call playbook was not retrievable for this article, so the guidance here is limited to what can be observed from the team’s own configuration: which alerts page, which alerts only notify, which escalation rules exist, and which alerts have been acknowledged without action in the past 90 days. The departing engineer should identify any alert that is known to be a false positive and any alert that has never fired but is expected to. The second category is the dangerous one.

Use structured prompts, not open-ended interviews

Open-ended knowledge transfer sessions tend to produce narrative, not procedure. A more effective format is a set of prompts tied to specific artifacts. For each backup repository, ask: what is the exact restore command, what is the expected duration, what is the failure mode if the repository is unavailable, and what is the last known-good snapshot. For each database cluster, ask: what is the stanza name, what is the retention policy, what is the archive command, and what is the last tested restore point. For each alert, ask: what does this alert mean, what is the first diagnostic step, and what is the escalation path if the first step does not resolve it.

These prompts should be answered in writing during the session, not after. The departing engineer should type the commands or paste the configuration while the remaining engineer watches. This produces a draft runbook that can be verified immediately. Google’s postmortem guidance emphasizes collaboration and knowledge-sharing at every stage, with real-time collaboration enabling rapid collection of data and ideas. The same principle applies to offboarding: the document should be built during the conversation, not reconstructed afterward.

For decision context, a useful prompt is: what would you do differently if you were designing this today, and why. This surfaces the tradeoffs that are not visible in the configuration. A pgBackRest stanza that uses a 30-day retention policy may be the result of a storage constraint that no longer exists, or a compliance requirement that still does. The departing engineer may be the only person who knows which. The answer should be recorded as a decision record, not as a runbook step, because it informs future changes rather than immediate operations.

Verify before access is revoked

The most common failure in offboarding knowledge extraction is that the remaining team believes it has captured the knowledge because it has written it down. The only reliable verification is execution by someone other than the departing engineer. For backup and restore, this means a restore drill during the two-week window, performed by the remaining engineer, with the departing engineer observing but not intervening. The drill should use the documented commands and the documented repository. If the restore fails, the failure is the finding, and the remaining time should be spent correcting the documentation and retesting.

For access and break-glass, verification means the remaining engineer uses the break-glass path to perform a low-risk action, such as listing a resource or reading a configuration value. If the path requires a credential that only the departing engineer holds, that is a finding. If the path requires a rotation step that is not documented, that is also a finding. The goal is not to test the credential’s power but to test the team’s ability to use it without the departing engineer.

For alerting and escalation, verification means the remaining engineer acknowledges a test alert and follows the documented escalation path. If the escalation path depends on a phone number or a schedule that only the departing engineer maintains, that dependency should be removed or reassigned before the last day. The verification should be recorded with a timestamp and the name of the person who performed it, so that the next offboarding has a baseline.

For runbooks and decision records, verification means the remaining engineer follows the runbook for a non-critical task, such as rotating a log file or checking a backup’s integrity. If the runbook requires a step that is not documented, the runbook is incomplete. The departing engineer should be available to answer questions during this verification, but the answers should be written into the runbook, not just spoken.

Integrate with access hygiene

Knowledge extraction and access revocation are not separate processes. The two-week window is also the window in which credentials should be rotated, break-glass paths should be reviewed, and least-privilege cleanup should be performed. The order matters: extraction should happen before revocation, but revocation should not wait until the last day. A useful pattern is to revoke non-essential access at the midpoint of the two weeks, after the backup and restore extraction is complete, and to revoke break-glass access only after the verification drill is complete.

This sequencing reduces the risk that the departing engineer’s access is used as a substitute for documentation. If the remaining engineer knows that the departing engineer can still restore the database, the incentive to verify the runbook is lower. Revoking access at the midpoint changes the incentive. It also surfaces dependencies earlier, when there is still time to resolve them.

Credential rotation should be treated as a verification step, not just a security step. If the departing engineer’s credentials are rotated and the remaining team can still perform all critical operations, the extraction is likely complete. If rotation breaks a critical operation, the extraction has a gap. The gap should be documented and resolved before the last day.

A lightweight process for a small team

A 2–15 person team does not need a dedicated offboarding program. It needs a repeatable sequence that fits in two weeks and produces artifacts that are useful after the departing engineer is gone. The sequence can be as simple as: day 1–3, backup and restore extraction with a restore drill; day 4–6, access and break-glass extraction with a break-glass drill; day 7–8, alerting and escalation extraction with a test alert; day 9–10, runbook and decision-record verification with a non-critical task. Each day should produce a written artifact and a verification record.

The artifacts should be stored where the remaining team already looks for operational information, not in a separate offboarding folder. A runbook that lives in a wiki page is more likely to be used than a runbook that lives in a shared drive. A decision record that lives next to the configuration it describes is more likely to be found than a decision record that lives in a meeting notes archive. The goal is to reduce the distance between the question and the answer.

The process should be reviewed after each offboarding. If a verification drill failed, the failure should be recorded and the process should be adjusted. If a runbook was incomplete, the gap should be noted and the template should be improved. Over time, the process becomes a form of institutional memory, not just a checklist. The Recovery Checklist Before You Need It is a useful companion for the backup and restore portion of this work, because it frames the drill as a pre-incident exercise rather than a post-departure scramble.

What cannot be captured in two weeks

Some knowledge resists extraction. Intuition about which alerts are noise, familiarity with the history of a service, and the social capital that comes from having worked with a vendor for years are not easily transferred. The goal of the two-week window is not to capture everything. It is to capture the knowledge that is both critical and verifiable: the restore command that works, the break-glass path that is tested, the escalation rule that is documented, and the decision record that explains why the configuration is the way it is.

The remaining knowledge should be treated as a known gap, not as a failure. A team that documents its gaps is more resilient than a team that assumes it has no gaps. The departing engineer should be asked to list the things they would want to know if they were staying, and that list should be added to the runbook as open questions. Some of those questions will be answered by the next incident. Some will be answered by the next hire. The important thing is that they are written down.

Frequently asked questions

How do we decide what to extract first when everything feels critical?

Start with the operations that have the longest recovery time and the least recent verification. A restore that has never been tested is higher priority than an alert that has been acknowledged without action for months. The blast radius of a failed restore is usually larger than the blast radius of a noisy alert.

What if the departing engineer is not available for the full two weeks?

Compress the schedule around the highest-blast-radius items. Backup and restore extraction can be done in a single day if the restore drill is scoped to a single repository and a single snapshot. Access and break-glass extraction can be done in half a day if the break-glass path is already documented. The verification steps are the ones that should not be skipped, because they are the only way to confirm that the extraction is accurate.

Should we record the extraction sessions?

Recording can be useful for reference, but it is not a substitute for written artifacts. A recording is difficult to search and easy to misinterpret. The written runbook and decision record should be the primary artifacts, with the recording as a backup. If the team uses recordings, they should be stored with the same access controls as the runbooks, because they may contain credential paths or internal hostnames.

How do we handle knowledge that is sensitive, such as break-glass credentials?

Sensitive knowledge should be documented in a way that describes the path without exposing the credential. The runbook should say where the credential is stored and how it is rotated, not what the credential is. The verification drill should confirm that the remaining engineer can access the credential through the documented path, not that the credential is written in the runbook.

What if the departing engineer is the only person who has ever performed a restore?

That is the highest-priority finding. The two-week window should be restructured to make the restore drill the first task, with the departing engineer observing and the remaining engineer executing. If the restore fails, the remaining time should be spent fixing the backup system, not documenting it. A documented restore that does not work is worse than an undocumented restore that does, because it creates false confidence.

How do we know when the extraction is complete?

The extraction is complete when the remaining team can perform all critical operations without the departing engineer’s access, and when the verification records show that those operations have been performed successfully. The absence of a departing engineer is not the test. The presence of a working runbook and a successful drill is the test.

Every new hire asks questions that the team has answered before. In a two-to-fifteen-engineer shop with no dedicated SRE function, those questions are the cheapest documentation audit you will ever run. The problem is that the answers usually live in a Slack thread, a pairing session, or someone’s memory, and the next hire asks the same thing six months later. A question log is a lightweight artifact that captures those questions, routes the ones that reveal real gaps into documentation, and leaves the rest alone.

This is not a knowledge-base project. It is a triage habit. The goal is to find the gaps before the next hire, not to document everything a new engineer might ever wonder about.

Why new-hire questions are a better audit than a documentation sprint

A documentation sprint asks the team to imagine what is missing. A new hire asks what is actually missing, in the order they encounter it. That order matters: it reflects the real path from laptop to production, not the org chart’s idea of how work flows.

The AWS Well-Architected Framework’s operational excellence pillar frames operations as something that should not be isolated from the teams it supports, and it recommends learning from operational events and sharing that learning. New-hire questions are an operational event with a predictable cadence. Treating them as a documentation input is consistent with that guidance, and it costs less than a dedicated review.

There is a second reason. In small teams, the people who answer questions are the same people who are on call. If the answer to “how do I rotate the age key for restic?” lives only in one engineer’s head, that engineer is a single point of failure for both documentation and incident response. Capturing the question and its answer is a small step toward reducing that dependency.

What belongs in the log

A question log is a list of questions, not a list of answers. The answer can be a link, a command, or a short paragraph. The question is the durable part because it tells you what the next person will need.

Useful fields:

  • Date asked. Helps you see whether a gap keeps recurring.
  • Question, in the asker’s words. Do not rewrite it into a documentation title yet. The raw phrasing is a signal about how people search.
  • Where it was asked. Slack channel, pairing session, pull request comment, or incident channel. This tells you where the answer already lives.
  • Answer or pointer. A link to the runbook, a command, or a note that no answer exists.
  • Gap? A yes/no judgment made during triage, not by the asker.
  • Action. One of: document, link from an existing doc, add to onboarding checklist, or no action.

The log can live in a shared document, a pinned Slack thread, or an issue tracker. The format matters less than the triage. A spreadsheet with a weekly review beats a beautiful wiki page that nobody reads.

Separating real gaps from normal onboarding friction

Not every question is a documentation gap. Some questions are the normal cost of learning a system, and documenting them would produce noise. The distinction is whether the answer is stable, discoverable, and needed by more than one person.

A question is probably a real gap when:

  • The answer exists but is not written down anywhere searchable.
  • The answer is written down but in a place the new hire could not reasonably find.
  • The answer requires a decision that is not recorded, such as why a particular backup tool was chosen.
  • Two different people give different answers.
  • The question is about a procedure that has operational consequences, such as restoring a database or rotating credentials.

A question is probably not a gap when:

  • It is about a one-off task that will not recur.
  • It is answered by reading the code, and the code is the source of truth.
  • It is a preference question with no operational consequence.
  • The answer is already in the onboarding checklist and the new hire simply had not reached that step.

The judgment is the hard part. The person doing triage should be someone who knows the system well enough to tell the difference, and who is not the new hire. In a small team, that is usually the on-call lead or the engineer who owns the relevant runbook.

How to run the log without adding process overhead

The failure mode of a question log is that it becomes a second job. The way to avoid that is to keep capture cheap and triage bounded.

Capture. Ask the new hire to add questions to the log as they arise, not after the fact. A single shared document with a table is enough. If the team already uses an issue tracker, a label like onboarding-question works. The important thing is that the new hire does not have to decide whether a question is worth logging. Log everything; triage later.

Triage. Once a week, or at the end of the new hire’s first month, spend fifteen minutes reviewing the log. For each question, mark it as a gap or not, and assign one action. The action should be small: add a paragraph to an existing runbook, add a link from the onboarding checklist, or open a ticket to write a new document. Do not try to write the document during triage.

Act. The action items go into the normal work queue. In a small team, that means they compete with feature work and incident follow-up. That is fine. The point is that they are visible and can be prioritized, not that they are done immediately.

Close the loop. When a gap is documented, add a link to the answer in the log. When the next hire asks the same question, the answer should be findable. If it is not, the documentation did not solve the problem.

What to document first

Not all gaps are equal. In a lean team running production without dedicated reliability coverage, the highest-value documentation is the kind that reduces the blast radius of an incident or the time to recover.

Prioritize questions about:

  • Backup and restore. How to restore a restic snapshot, how to verify a pgBackRest backup, how to find the right WAL segment for point-in-time recovery. The PostgreSQL documentation on continuous archiving and point-in-time recovery is explicit that you need a continuous sequence of archived WAL files extending back at least as far as the start time of your base backup, and that you should set up and test your archiving procedure before taking your first base backup. A new hire asking “how do I know the WAL archive is working?” is pointing at a real operational gap.
  • Access and break-glass. How to get production access, how to use the break-glass path, and what to do afterward. These questions often reveal that the process is undocumented or that the documentation is stale.
  • Incident response. Who to page, how to escalate, where the incident channel is, and what the postmortem process looks like. PagerDuty’s public incident response documentation is used to prepare new employees for on-call responsibilities, which is a useful model for what a new hire needs to know before their first shift.
  • Monitoring and alerts. What the alerts mean, which ones are actionable, and which ones are noise. A new hire asking “why did this alert fire?” may be pointing at an alert that needs tuning.

Questions about local development setup, code style, and tool preferences are lower priority unless they recur across hires.

Using the log to improve onboarding, not just documentation

The log has a second use: it tells you what the onboarding checklist is missing. If three consecutive hires ask how to request database access, the checklist is wrong, not the documentation.

Review the log at the end of each onboarding cycle and update the checklist. This is a small change that compounds. It also gives the new hire a sense that their questions are improving the team’s process, which is consistent with the blameless, learning-oriented culture described in Google’s SRE book. The book notes that postmortems are expected after significant undesirable events, and that the goal is to understand contributing causes without indicting individuals. The same posture applies to onboarding questions: the question is not a failure by the new hire, it is a signal about the system.

Metrics that tell you the log is working

You do not need a dashboard. A few simple counts are enough:

  • Repeat questions. How many questions in this onboarding cycle were already in the log from a previous cycle? If the number is going down, documentation is working. If it is flat or going up, the answers are not discoverable.
  • Time to first independent change. How long before the new hire shipped a change without asking for help? This is a rough measure, but it is the outcome the log is meant to improve.
  • Gap closure rate. How many gaps marked during triage were actually documented? A low rate means the actions are not being prioritized.
  • Questions about incidents. How many questions are about how to respond to a specific failure mode? These are the highest-value gaps because they affect recovery time.

None of these are precise. They are directional. The point is to notice when the log is filling up with the same questions and nothing is changing.

A worked example

Suppose a new hire joins a team running PostgreSQL on a single primary with pgBackRest and restic for file-level backups. In their first two weeks, they ask:

  1. How do I restore a single table from a pgBackRest backup?
  2. Where are the restic repository credentials stored?
  3. How do I know the WAL archive is current?
  4. What is the difference between the staging and production alert channels?
  5. How do I request access to the production database?

During triage, the team marks questions 1, 3, and 5 as gaps. Question 2 is answered by a link to the secrets manager documentation, which the new hire had not found. Question 4 is answered by a link to the monitoring runbook, which already existed but was not linked from the onboarding checklist.

The actions are: add a restore example to the pgBackRest runbook, add a WAL archive check to the daily operations checklist, and add the access request process to the onboarding checklist. None of these are large. Together they close three gaps that would otherwise have been rediscovered by the next hire.

What not to do

Do not turn the log into a documentation backlog that grows faster than it shrinks. If a question is marked as a gap and not documented within a reasonable time, either the gap is not important or the team does not have capacity. Both are useful signals. A backlog of fifty undocumented gaps is worse than a log of five that get closed.

Do not ask the new hire to write the documentation. They do not yet know enough to write it accurately, and the exercise can feel like busywork. The new hire’s job is to ask questions and log them. The team’s job is to answer and document.

Do not document everything. The goal is to reduce repeat questions and improve incident response, not to produce a complete manual. A lean team’s documentation should be a map of the terrain that matters, not a transcript of every conversation.

Where this fits with other resilience practices

A question log is not a substitute for a recovery drill or a postmortem. It is a complement. The recovery checklist you write before you need it, the postmortem you write after an incident, and the question log you keep during onboarding all serve the same purpose: they move knowledge from one person’s head into a form the team can use.

If you are starting from nothing, start with the recovery checklist. It is the highest-stakes document a lean team can have. Then add the question log. It is the cheapest way to find the next gap.

FAQ

How long should we keep the question log?

Keep it as long as it is useful. A log that spans several onboarding cycles is more valuable than one that is reset each time, because it shows which questions recur. If the log becomes too large to review in fifteen minutes, archive the closed items and keep the open gaps.

Who should own triage?

The person who owns the relevant runbook or the on-call lead. In a small team, that is often the same person. The owner should not be the new hire, because the judgment about what is a real gap requires system knowledge.

What if the team does not have time to document the gaps?

Then the log is telling you something about capacity. Either the gaps are not important enough to prioritize, or the team is too stretched to improve its own documentation. Both are worth knowing. The log does not create the problem; it makes it visible.

Should the log be public to the whole team?

Yes, unless it contains sensitive information. Questions about access paths and break-glass procedures may need to be restricted. In that case, keep the question in the log and put the answer in the restricted runbook.

How is this different from a runbook?

A runbook is a procedure. A question log is a list of things people did not know. The log feeds the runbook. It is the input, not the output.

Offboarding is a reliability event disguised as an HR event. When an engineer leaves a 2–15 person team, the same person who held production access, carried a pager, and knew which runbook to open at 03:00 is now a set of credentials that must be revoked. The failure mode is not that revocation is forgotten; it is that revocation is done in the wrong order, and the on-call rotation loses a responder, a break-glass path, or a backup credential at the moment it is needed.

This checklist is written for teams that are their own SRE function. It assumes AWS, GCP, or bare metal, PostgreSQL with WAL archiving, restic or pgBackRest for backups, and PagerDuty or an equivalent escalation tool. It is a sequence, not a policy document. The goal is to remove a person’s access within 48 hours while keeping the rotation intact.

Why 48 hours, and why order matters

NIST SP 800-53 Rev. 5, the control catalog for federal information systems, includes an Access Control family covering account management. The operational reading for a lean team is that revocation must be a defined procedure with an owner, not an ad hoc action. The 48-hour window is a team-level target, not a NIST requirement. Treat it as a service-level objective for your own process.

Order matters because access is layered. A departing engineer may hold a long-lived IAM user with an access key, a federated identity through an IdP, a PagerDuty user with a schedule override, a database role with a password in a password manager, and an SSH key on a bastion. Revoking the IAM user first can break a script that the rotation depends on. Revoking the PagerDuty user first can leave a gap in the escalation policy. The sequence below removes human access first, then service access, then shared secrets, then the pager identity.

Phase 1: Freeze and inventory (0–4 hours)

Before removing anything, capture what exists. This is the step that prevents the 03:00 surprise.

  • List all IAM users, roles, and access keys associated with the person in AWS. In AWS, temporary credentials issued through AWS STS are short-lived and expire on their own; long-lived IAM user access keys do not. The AWS IAM documentation states that temporary credentials can be configured to last from a few minutes to several hours and that after they expire, AWS no longer recognizes them. If the departing engineer’s access is federated through an IdP, the long-lived key inventory may be smaller than expected.
  • List all GCP principals. Google Cloud IAM grants roles to principals on resources, and allow policies are inherited down the resource hierarchy. A role granted at the organization or folder level applies to every descendant resource. Revoking a project-level binding does not remove an inherited organization-level binding. The Google Cloud IAM overview documents this inheritance explicitly: to understand who can access a resource, you must view the resource’s allow policy and its ancestors’ allow policies.
  • List PagerDuty schedule memberships, escalation policy steps, and any user-level notification rules. The PagerDuty integrations documentation describes the platform as connecting tools and data to incident management; the operational point is that a user is a first-class object in schedules and escalation policies, not just an email address.
  • List database roles and authentication methods. For PostgreSQL, check pg_hba.conf for any trust or md5 entries tied to a specific host or user, and check for roles with SUPERUSER or REPLICATION attributes.
  • List backup credentials. If restic repositories are accessed with a password stored in a shared vault, note who knows it. If pgBackRest uses an SSH key or a cloud identity, note which one.

Write the inventory to a ticket or a decision record. The inventory is the artifact that makes the rest of the checklist auditable.

Phase 2: Remove human access (4–24 hours)

Human access is the interactive path: console logins, CLI sessions, SSH sessions, and database sessions. Remove it before touching service accounts.

AWS

If the engineer has an IAM user, deactivate the access keys and delete the login profile. If the engineer accesses AWS through a federated identity, disable the identity in the IdP. The AWS documentation notes that temporary credentials are the basis for roles and identity federation and that they do not have to be explicitly revoked when no longer needed because they expire. That property is useful during offboarding: a federated session that is already issued will expire on its own, but the IdP must be disabled so new sessions cannot be issued.

Check for any IAM roles with a trust policy that allows the departing engineer’s principal to assume them. A role trust policy is not a permission grant to the person, but it is an access path. Remove the principal from the trust policy if the role was created for that person.

GCP

Remove the principal from all allow policies where it appears. Because of policy inheritance, check the organization, folder, and project levels. The Google Cloud IAM overview notes that the IAM API is eventually consistent: a write followed immediately by a read may return an older version, and changes may take time to affect access checks. Do not assume a revocation is effective the moment the API returns success. Verify with a read after a short delay, or use the policy troubleshooter.

If the engineer has a service account key, delete the key. Service account keys are long-lived credentials and are a common offboarding gap because they are often created for automation and then used interactively.

Bare metal and SSH

Remove the public key from authorized_keys on every host the engineer could reach. If you use a configuration management tool, remove the key from the source of truth and let the tool converge. If you use a bastion with short-lived certificates, revoke the certificate authority’s trust for that principal if the CA supports it; otherwise, wait for the certificate to expire and confirm the expiry window.

PostgreSQL

Revoke login on the role: ALTER ROLE username NOLOGIN;. Do not drop the role immediately if it owns objects or has granted privileges; dropping a role that owns objects requires reassigning ownership first. Check pg_shdepend for dependencies. If the role has SUPERUSER, revoke it before disabling login so the change is visible in the audit log.

If the engineer had a .pgpass file or a password in a shared vault, rotate the password for any role that was shared. A shared role is a single point of failure during offboarding; if two engineers used the same database role, disabling one person’s login does not remove the other person’s access, but it also does not remove the departing engineer’s access if the password is still known.

Phase 3: Service accounts and automation (24–36 hours)

Service accounts are where offboarding breaks the rotation. A departing engineer may have created a service account for a synthetic check, a backup job, or a deployment pipeline. If that service account is disabled without replacement, the check stops running and the rotation loses a signal.

For each service account the engineer owned or used:

  • Identify what depends on it. Check CI/CD pipelines, cron jobs, Kubernetes service accounts, and monitoring integrations.
  • Create a replacement service account with the minimum permissions needed, or transfer ownership to a shared team identity.
  • Update the dependency to use the replacement.
  • Disable the old service account and delete its keys.

In GCP, service accounts are principals. The IAM overview distinguishes human users from workloads and notes that service accounts are workload principals. That distinction is useful: a service account should not be used interactively, and an interactive user should not be the sole owner of a service account.

In AWS, the equivalent is an IAM role for EC2 or a role assumed by a pipeline. The AWS documentation describes roles for Amazon EC2 as a way to provide temporary credentials to instances without storing long-term credentials. If the departing engineer created an IAM user with an access key for a script, replace it with a role.

Phase 4: Backups and recovery paths (36–44 hours)

Backup access is the most commonly missed offboarding item because it is often shared. The question is not whether the departing engineer can restore; it is whether the remaining team can restore without them.

PostgreSQL WAL archiving

PostgreSQL’s continuous archiving documentation describes WAL archiving as a sequence of segment files copied by an archive_command or an archive_library. The archive command runs under the ownership of the PostgreSQL server user. If the archive command depends on a credential that the departing engineer controlled, the archive will fail after offboarding. Check the command for references to a specific home directory, an SSH key, or a cloud credential file.

The documentation also notes that if the archive command fails repeatedly, the pg_wal/ directory will fill with segment files, and if the file system fills, PostgreSQL will perform a PANIC shutdown. That is a production outage caused by an offboarding gap. Verify that the archive command uses a credential owned by the service, not by a person.

Test the restore path. A backup that has never been restored is a hypothesis. The Gray Haven article Write the Recovery Checklist Before You Need It describes the practice of writing the recovery procedure before an incident; use it here to confirm that the remaining team can perform a point-in-time recovery without the departing engineer’s credentials.

restic and pgBackRest

If restic repositories are encrypted with a password that the departing engineer knows, rotate the repository password. restic supports multiple keys per repository; add a new key for the team, then remove the old key. If pgBackRest uses a repository host with SSH keys, replace the key and update the configuration.

If the backup repository is in cloud storage, check the bucket or container policy for any principal that matches the departing engineer. In GCP, a bucket policy is an allow policy on the bucket resource; inherited policies from the project may also apply. In AWS, check the bucket policy and any IAM role that grants access.

Phase 5: PagerDuty and the on-call rotation (44–48 hours)

The pager identity is the last thing to remove because it is the thing that keeps the rotation working. Removing a user from PagerDuty before updating the schedule can create a gap; updating the schedule before removing the user can leave a stale notification path.

Sequence:

  1. Add a replacement responder to every schedule the departing engineer was on. If the team is small, this may mean the remaining engineers absorb the shift. Confirm that the replacement has the same escalation policy coverage.
  2. Update escalation policies to remove the departing engineer as a step. If the engineer was the only person in a step, replace the step before removing the user.
  3. Check for schedule overrides. A departing engineer may have an override that extends beyond their last day.
  4. Remove the user from PagerDuty. If the user is also the account owner or billing contact, transfer those roles first.
  5. Verify that a test incident routes to the remaining responders. PagerDuty’s integrations documentation describes email and API integrations; use a test event through the same path your production alerts use.

If the departing engineer was the only person who knew how to acknowledge a specific alert, that is a runbook gap, not a PagerDuty gap. Fix the runbook before the last day.

Break-glass access

Break-glass access is the credential you use when normal access is unavailable. It is also the credential most likely to be tied to a person. If the departing engineer held the break-glass credential, rotate it. If the break-glass credential is a shared password in a vault, change the password and update the vault entry.

The CISA Zero Trust Maturity Model page was not retrievable at the time of writing, so this article does not cite it. The operational principle is narrower: break-glass access should be auditable and time-bound. If your break-glass path is a long-lived IAM user with an access key, replace it with a role that requires MFA and issues temporary credentials. AWS documents temporary credentials as short-term and expiring; that property limits the window in which a lost break-glass credential can be used.

What to verify before closing the ticket

  • The departing engineer cannot authenticate to AWS, GCP, or any bare-metal host.
  • No service account or automation job depends on a credential the engineer controlled.
  • WAL archiving is still running and the archive command does not reference a personal credential.
  • A restore drill has been performed by a remaining team member within the last quarter.
  • The PagerDuty schedule has no gaps and a test incident routes correctly.
  • Break-glass credentials have been rotated.
  • The inventory and the actions taken are recorded in a decision record or ticket.

FAQ

Do we need to revoke temporary credentials explicitly?

No. AWS documents temporary credentials as expiring after a configured lifetime, after which AWS no longer recognizes them. The action required during offboarding is to prevent new temporary credentials from being issued, which means disabling the identity in the IdP or removing the principal from the role trust policy.

What if the departing engineer is the only person who knows the backup password?

That is a single point of failure that predates the offboarding. Rotate the password before the last day, add a new key for the team, and verify a restore. If the password is lost, the backup may be unrecoverable; treat that as an incident and document it.

How do we handle a shared database role?

Replace it with individual roles. A shared role cannot be revoked for one person without affecting others. If individual roles are not feasible, rotate the password and distribute it only to the remaining team, then plan to split the role.

What if the engineer leaves before we finish the checklist?

Prioritize human access removal and PagerDuty coverage. Service accounts and backup credentials can be rotated after the last day, but an active human credential and an uncovered pager are immediate risks. If the engineer is cooperative, ask them to transfer ownership of service accounts before their last day.

Does NIST SP 800-53 require a specific revocation timeframe?

No. SP 800-53 Rev. 5 includes an Access Control family covering account management, but the specific requirements and any timeframe are organizational decisions. The 48-hour window in this article is a team-level target, not a NIST requirement.

Sources

Every lean team has one. A crontab on a box that predates the current on-call rotation, running a job that touches production data, owned by a user account nobody remembers creating. It works, so it stays. Then it fires at 03:00 on a Sunday, the person it pages has never seen the script, and the first fifteen minutes of the incident are spent answering a more basic question: what is this job, and who was supposed to know about it?

This is not a story about cron being fragile. Cron is remarkably durable. The problem is that cron encodes almost no ownership metadata, and the metadata it does encode is easy to misread. An audit of scheduled jobs is really an exercise in reconstructing intent from a format that was never designed to carry it.

What the crontab format actually tells you

The crontab(5) format is five time-and-date fields, an optional username in system crontabs, and a command. That is the whole schema. There is no owner field, no description field, no ticket reference, no expiry. The manual page is explicit that commands in a given crontab execute under the user who owns that crontab, and that LOGNAME and HOME are set from the owner’s /etc/passwd entry while SHELL defaults to /bin/sh unless overridden. Ownership, in the operational sense, is inferred from the account, not declared.

That inference breaks in predictable ways. A job in /etc/cron.d/ carries a username field because those are system jobs used for more than one user. A job in a personal crontab does not, because the account is the identity. When the account is a shared service account like deploy or app, the crontab tells you which Unix user runs the job and nothing about which human is accountable for it. Those are different questions, and conflating them is how a job ends up firing at the wrong person.

Two format details matter more than they look. First, comments are not allowed on the same line as a cron command; the manual page notes they are treated as part of the command. So the tempting habit of appending # owner: alice to a job line does not document anything — it changes the command. Second, if both day-of-month and day-of-week are restricted, the job runs when either matches. A line like 30 4 1,15 * 5 runs at 04:30 on the 1st, the 15th, and every Friday. Teams that read it as “the 1st and 15th, if it’s a Friday” have a job running roughly four times more often than they believe.

Time handling is the other quiet source of surprise. The manual page states that non-existent times during a daylight-saving transition never match, so jobs scheduled in the skipped hour do not run, and times that occur twice cause matching jobs to run twice. The cron(8) page describes the daemon’s own handling: for local time changes under three hours, jobs that would have run in a skipped interval are run immediately, and a backward adjustment avoids running the same job twice. These two descriptions are not contradictory — they describe different layers — but the practical upshot for an audit is that “it runs at 02:30 daily” is not a stable statement across DST boundaries. If a job’s correctness depends on running exactly once per day, that assumption deserves a test, not a shrug.

Where the job’s output goes, and why that determines who gets paged

Failure attribution for cron jobs is mostly a question of where output and errors land. The cron(8) page is direct: any output is mailed to the owner of the crontab, or to the address in MAILTO if set. The crontab(5) page adds the detail that MAILTO="" suppresses mail entirely, and that MAILFROM controls the envelope sender. It also notes that -s directs job output to syslog, which is useful when sendmail is absent or mail is disabled.

This is the mechanism behind a common failure mode. A job’s MAILTO points at a distribution list that was retired, or at an individual who left, or at nothing because someone set MAILTO="" during a noisy week and never reverted it. The job fails silently for months. The first signal is a downstream symptom — stale data, a missing report, a backup that was never verified — and the person who notices is whoever depends on the output, not whoever owns the job.

An audit should therefore record, for each job: the effective MAILTO, whether mail is suppressed, whether -s is in use, and where the job’s own logging goes if it writes any. A job that logs to a file nobody reads and mails to an address nobody monitors is, for incident purposes, unowned regardless of what the crontab says.

System cron versus systemd timers: the logging and dependency difference

Lean teams often inherit a mix. System cron and systemd timers differ in ways that affect an audit’s scope.

Cron’s model is stateless and time-based. The daemon examines stored crontabs each minute and runs what matches. It reloads changed crontabs without a restart, either by checking modtimes every minute or via inotify. It has no concept of a job depending on another job, no built-in retry, and no structured exit-status record beyond whatever the command itself emits. The cron(8) page documents clustering support via a shared /var/spool/cron and a .cron.hostname file, which is worth knowing if you have ever wondered why a job runs on exactly one host in a group — and worth flagging, because that file is itself an undocumented dependency.

Systemd timers, by contrast, are units. They carry a description, they can declare dependencies and ordering, they record activation results in the journal, and they can be inspected with systemctl list-timers. For an audit, the practical difference is that a timer’s failure is queryable after the fact through journald, while a cron job’s failure may exist only as a mail message that was never delivered or a syslog line that rotated away.

This is a tradeoff, not a verdict. Cron is simpler to reason about and present on nearly everything. Timers give you structured state and dependency handling at the cost of more configuration surface and a unit-file review that most teams never do. The audit question is not which is better; it is whether you can answer “did this job run, and did it succeed?” for every scheduled job you have. If the answer depends on reading mail, the answer is no.

Backup jobs deserve their own audit pass

Scheduled jobs that touch backups are the ones where an undocumented crontab entry has the highest cost, because a backup job that has been failing quietly is indistinguishable from a backup job that has been succeeding quietly until the day you need to restore.

PostgreSQL’s continuous archiving documentation is explicit about the dependency chain. To recover using continuous archiving, you need a continuous sequence of archived WAL files extending back at least as far as the start time of your base backup. The archive command must return zero exit status if and only if it succeeds; a nonzero status tells PostgreSQL the file was not archived and it will retry periodically. The documentation also warns that when the archive command is terminated by a signal or fails with an exit status greater than 125, the archiver process aborts and restarts, and that failure is not reported in pg_stat_archiver. That last point is the audit-relevant one: a monitoring check built only on pg_stat_archiver can miss a class of archiving failures entirely.

The documentation further notes that if archiving falls significantly behind, the amount of data that would be lost in a disaster increases, and pg_wal/ accumulates unarchived segments that can eventually exhaust disk. It recommends monitoring the archiving process. For a lean team, that recommendation translates into a specific audit item: for every scheduled job that archives WAL, record what monitors it, what the alert threshold is, and whether the monitor distinguishes “archiving is behind” from “archiving is broken.”

pgBackRest’s user guide is useful here for a different reason: it documents the retention and dependency semantics that determine whether a restore will actually work. A differential backup requires the prior full backup to be valid; an incremental backup requires all prior incrementals back to the prior differential, the differential itself, and the full backup. The guide states that a restore requires the backup files and one or more WAL segments. The audit implication is that a retention policy which expires a full backup while keeping incrementals that depend on it produces a repository that looks populated and cannot be restored. If your scheduled expire job runs on a crontab nobody owns, that is the failure you are auditing for.

Restic’s documentation covers the same territory from the repository side. It documents restic check for integrity and consistency verification, and it documents exit status codes and JSON output for scripting. The relevant audit point is that a scheduled restic check that runs but whose output goes to a mailbox nobody reads provides no assurance. The check has to be wired to something that fails loudly.

None of this substitutes for an actual restore. A scheduled backup job that has never been restored from is a hypothesis. The recovery checklist approach — writing down the steps before you need them — is the natural companion to this audit, because the audit tells you which jobs exist and the checklist tells you what to do when one of them turns out to have been lying.

A workable audit procedure

The goal is a single inventory that answers, for every scheduled job: what runs, as whom, on which hosts, writing where, alerting whom, and who is accountable. The procedure below is deliberately mechanical.

1. Enumerate. Collect every source of scheduled execution: user crontabs in /var/spool/cron, system crontabs in /etc/crontab and /etc/cron.d/, /etc/anacrontab, systemd timers via systemctl list-timers --all, and any scheduler that lives inside an application (Celery beat, Sidekiq cron, Kubernetes CronJobs). The last category is the one most often missed, because it is not on the box you are auditing.

2. Resolve the effective environment. For each cron entry, record the effective SHELL, HOME, PATH, MAILTO, and CRON_TZ. The manual page notes that PATH is set by cron unless -P is used, and that HOME and SHELL can be overridden in the crontab while LOGNAME cannot. A job that works interactively and fails under cron is usually a PATH or HOME difference, and recording the effective values turns that from a debugging session into a lookup.

3. Trace the command. Follow the command to its script, and record what the script does at a level a new on-call engineer could act on. If the script calls other scripts, follow those too. The audit output is not a copy of the code; it is a sentence or two of intent plus a pointer to the source.

4. Identify the accountable human. This is the step that requires a conversation, not a command. For each job, name a person, not a team. If nobody will claim it, that is the finding — not a gap in the spreadsheet. A job with no accountable owner is a candidate for deletion, and deletion is a legitimate audit outcome.

5. Record the failure path. For each job, answer: if this fails, who finds out, how, and how long would it take? If the answer is “nobody” or “when someone notices the data is wrong,” write that down. It is the most important column in the inventory.

6. Test the schedule assumptions. For any job whose correctness depends on running exactly once per interval, verify the DST behavior and the day-of-month/day-of-week semantics against the manual page rather than against intuition. For jobs using step values, remember that steps are evaluated within their field: */35 in the minute field runs at minute 0 and minute 35, not every 35 minutes.

7. Validate syntax before changes. The crontab(5) page notes that crontab syntax can be tested before installation using the -T option. Use it. An audit that introduces a broken crontab is a self-inflicted incident.

What to do with the inventory

The inventory is not the deliverable. The deliverable is a set of decisions: jobs to delete, jobs to move to a timer with structured logging, jobs to add monitoring for, jobs to document in a runbook, and jobs to hand to a named owner.

Two of those decisions deserve emphasis for lean teams.

Deletion is underrated. A scheduled job that has run without anyone noticing its output for a year is, in most cases, not load-bearing. The cost of keeping it is not zero: it is another thing that can page someone at 03:00, another thing to audit next year, another line in the inventory that dilutes attention. The audit is the moment to ask whether the job still earns its place.

Monitoring is the other. A scheduled job without a failure signal is a job you will learn about from its consequences. The minimum viable signal is a heartbeat: something that alerts when the job doesn’t run, not just when it exits nonzero. Cron’s mail-based failure reporting covers the second case and not the first, which is why a job that never starts is the quietest failure of all.

Postmortems for scheduled-job failures

When a scheduled job does cause an incident, the postmortem is where the audit’s findings become durable. Google’s SRE book chapter on postmortem culture is a useful reference for the practice, and its framing applies directly here. The chapter defines a postmortem as a written record of an incident, its impact, the actions taken to mitigate or resolve it, the root cause or causes, and the follow-up actions to prevent recurrence. It lists common triggers including user-visible downtime beyond a threshold, data loss of any kind, on-call engineer intervention, resolution time above a threshold, and monitoring failure — the last of which is exactly what an unmonitored cron job produces.

The chapter’s argument for blamelessness is the part that matters for this topic. It states that a blameless postmortem focuses on identifying contributing causes without indicting any individual or team, and that it assumes everyone involved acted with good intentions given the information they had. The reasoning is practical: if blame is the norm, people stop surfacing issues, and the issues that go unsurfaced are the ones that cause the next incident. For a scheduled job that fired at the wrong person, the blameless framing points at the real question — why did the system allow a job to exist without a named owner and a failure signal? — rather than at whoever wrote the crontab entry three years ago.

The chapter also recommends defining postmortem criteria before an incident occurs, so that everyone knows when one is required. For lean teams, a workable criterion is: any scheduled job whose failure was discovered by a human noticing a downstream symptom gets a postmortem. That is a narrow, testable trigger, and it captures the class of failure this audit is designed to reduce.

Keeping the inventory alive

An audit that produces a document and no process is a snapshot that decays. The inventory needs a home and a review cadence. Two practices keep it current without much overhead.

First, tie the inventory to onboarding and offboarding. When someone joins, the jobs they inherit are part of their onboarding. When someone leaves, the jobs they owned are part of their offboarding checklist. This is the same access-hygiene discipline applied to scheduled work, and it is the only reliable way to prevent the “owned by a departed engineer” failure mode.

Second, make the inventory a required input to any change that adds a scheduled job. A new cron entry or timer should not merge without an owner, a failure signal, and a line in the inventory. This is a small process cost paid at the moment of creation, when the context is fresh, instead of a large reconstruction cost paid during an incident.

Questions that come up during these audits

How do I find crontabs I don’t know about? Start with the documented locations: /var/spool/cron for user crontabs, /etc/crontab and /etc/cron.d/ for system jobs, and /etc/anacrontab. Then check systemd timers and any application-level scheduler. The cron(8) page lists the files and directories the daemon checks, which is the authoritative starting point.

Why did a job run twice? Check the DST behavior first. The crontab(5) page notes that times occurring more than once during a transition cause matching jobs to run twice. Also check whether the job exists in more than one crontab, or on more than one host in a group where clustering is not configured.

Why didn’t a job run at all? Check whether the scheduled time fell in a DST gap, which the manual page says will never match. Check whether the crontab file is missing a trailing newline, which the manual page warns causes cron to consider the crontab at least partially broken and write a warning to syslog. Check whether the file permissions satisfy cron’s requirement that crontab files be regular files or symlinks to regular files, not executable or writable by anyone but the owner.

Where did the job’s output go? Check MAILTO in the crontab. If it is set and non-empty, mail goes there. If it is set but empty, no mail is sent. If it is unset, mail goes to the crontab owner. If -s is in use, output goes to syslog instead.

How do I know a backup job actually produced a restorable backup? You don’t, from the job’s exit status alone. PostgreSQL’s documentation requires a continuous WAL sequence back to the base backup’s start time, and pgBackRest’s documentation describes the dependency chain between full, differential, and incremental backups. The only reliable verification is a restore, which is why the recovery checklist and the audit belong together.

Should we migrate everything to systemd timers? Not necessarily. Timers give you structured logging and dependency handling; cron is simpler and more portable. The decision should follow from whether you can currently answer “did this job run and succeed?” for each job. If you can, the migration is optional. If you can’t, the migration is one way to fix that, and adding a heartbeat monitor is another.

The point of the exercise

An unowned crontab entry is not a crisis. It is a small, accumulating liability that stays invisible until the moment it becomes expensive. The audit is cheap: a few hours of enumeration and a conversation per job. The alternative is discovering the inventory during an incident, one job at a time, with the clock running and the person who wrote the job unavailable.

The output that matters is not the spreadsheet. It is the set of jobs you deleted, the jobs you gave an owner, and the jobs you wired to a signal that fires before a human notices the consequences. Everything else is documentation, and documentation that nobody reads is just another unowned artifact.

A five-person team with no dedicated SRE function usually shares one database admin password. It lives in a password manager entry, a CI secret, and someone’s shell history. The credential is convenient and it is also the reason a single offboarding takes a week, a leaked log line becomes a full-database incident, and nobody can answer “who ran that migration” without guessing. This is a field note on replacing that shared credential with individual, short-lived Vault-issued logins over a weekend, and on the parts that realistically need a staged rollout.

What the weekend can actually deliver

The realistic weekend scope is: enable Vault’s database secrets engine against PostgreSQL, define one or two roles that map to the privileges the team already uses, wire human authentication through an auth method, and cut over interactive access. Automation cutover, break-glass design, and audit review are usually the following week. Trying to do all of it in 48 hours is how teams end up with a half-migrated credential and two sources of truth.

Vault’s PostgreSQL database secrets engine generates credentials dynamically based on configured roles, and it also supports static roles. The plugin is named postgresql-database-plugin, and the documented setup is a vault secrets enable database followed by a vault write database/config/... with a connection URL and a privileged username. The role definition carries the SQL that Vault executes to create the credential, for example CREATE ROLE "{{name}}" WITH LOGIN PASSWORD '{{password}}' VALID UNTIL '{{expiration}}'; plus whatever grants the role needs. Default and max TTL are set per role; the documentation example uses default_ttl="1h" and max_ttl="24h". That is the whole mechanism. It is not exotic, and it does not require a platform team.

Two details from the same documentation matter for a small team. First, the plugin supports password_authentication="scram-sha-256", which is the setting you want on any PostgreSQL 14+ instance rather than the older md5 default. Second, the rootless static-role workflow is Vault Enterprise only, and it explicitly does not support dynamic roles. If you are on Vault Community, plan on a single privileged connection user for the database secrets engine and treat that user as a high-value credential in its own right.

Human authentication: keep it boring

For five engineers, the userpass auth method is sufficient and auditable. It allows users to authenticate with a username and password configured directly on the auth method, and it lowercases all submitted usernames, so Mary and mary are the same entry. The documented enablement is vault auth enable userpass, then vault write auth/userpass/users/<name> password=... policies=.... User lockout is enabled by default from Vault 1.13 onward, with a default threshold of 5 failed attempts, a 15-minute lockout duration, and a 15-minute counter reset. That default is a reasonable starting point; it also means a fat-fingered engineer during an incident will be locked out for 15 minutes, so decide in advance whether the on-call policy needs a second path.

If your team already has an identity provider, an OIDC or LDAP auth method is a better long-term choice than userpass because offboarding happens in the IdP rather than in Vault. The weekend tradeoff is that userpass is one command per person and OIDC is a configuration project. Either is defensible; the failure mode to avoid is leaving the shared password in place “until we finish the OIDC work.”

PostgreSQL roles: separate the human from the machine

The mistake that makes this migration painful is conflating the human login with the application login. PostgreSQL’s CREATE USER is an alias for CREATE ROLE with LOGIN assumed by default, and the privilege model is role-based: a role holds privileges directly, through role membership, and through anything granted to PUBLIC. That means you can build a small number of privilege-bearing roles and hand out membership, rather than granting table privileges to every human individually.

A workable shape for a five-person team:

  • app_rw — the role the application connects as. Owned by the application’s service account, not by humans.
  • oncall_read — CONNECT on the database, USAGE on the relevant schemas, SELECT on the tables the team actually needs to inspect during an incident.
  • oncall_write — the narrower set of write privileges the team has historically needed for hotfixes, granted explicitly rather than via ALL PRIVILEGES.
  • migration — the role that owns schema changes, used by the migration tool and by humans only during a planned change window.

Vault roles then map to these: a readonly role whose creation statement grants oncall_read, and a breakglass role whose creation statement grants oncall_write with a short TTL. The GRANT documentation is explicit that granting a privilege on a table does not extend to sequences used by that table, including sequences tied to SERIAL columns, so any role that inserts into a table with a serial primary key needs USAGE on the sequence as well. This is the kind of detail that turns a “read-only” role into a broken role at 02:00.

One more note from the GRANT documentation worth internalizing: superusers can access all objects regardless of object privilege settings, and the documentation compares this to root on a Unix system and advises against operating as a superuser except when absolutely necessary. The shared admin password was almost certainly a superuser. The point of this migration is to make superuser access rare, named, and time-bound.

Break-glass without reintroducing the shared password

The objection that stalls this migration is always the same: “what happens when Vault is down and we need the database?” The answer is not a shared password in a sealed envelope. It is a separate, deliberately awkward path that is still individual and still logged.

OWASP’s Secrets Management Cheat Sheet has a section on downtime, break-glass, backup and restore, and its general guidance is that secrets should be centralized, access-controlled on a least-privilege basis, and audited. The cheat sheet’s auditing section lists what should be captured at minimum: who requested a secret and for what system and role, whether the request was approved or rejected, when the secret was used and by whom, when it expired, whether there were attempts to reuse expired secrets, authentication and authorization errors, and administrative actions. That list is a good specification for what your break-glass path must produce.

A design that satisfies it without a shared password:

  1. A dedicated PostgreSQL role, breakglass, with LOGIN and a password stored in a separate secrets manager that is not Vault — for example, the cloud provider’s native secret store, or a hardware-backed password manager entry that requires a second approver to reveal.
  2. The password is rotated after every use, and the rotation is a documented runbook step, not a memory.
  3. Every use is announced in the incident channel before the credential is revealed, and the reveal is logged by the secrets manager.
  4. PostgreSQL’s own connection logging captures the login, so the database side has an independent record.

The tradeoff is real: this path is slower than typing a shared password. That is the point. A break-glass path that is as convenient as normal access will become normal access.

What to stage rather than rush

Interactive human access is the safe first cutover because a broken login is immediately visible and the blast radius is one engineer. The following should be staged:

  • Application credentials. Moving an application from a static password to Vault dynamic credentials requires the application to fetch and refresh credentials, or a sidecar to do it. OWASP describes the sidecar pattern — a Vault Agent sidecar authenticating with a Kubernetes service account, writing the secret to a shared in-memory volume, and refreshing it periodically — as a way to decouple the application from the secrets manager. That is a code and deployment change, not a weekend change.
  • CI and migration tooling. These often hold the longest-lived credentials and are the hardest to rotate because a failed rotation breaks deploys. Give them their own Vault role with a TTL that matches the deploy cadence, and test the rotation in a staging environment first.
  • Monitoring and alerting. If your synthetic checks or backup jobs connect to PostgreSQL, they need their own role. A backup job that fails silently because its credential expired is a worse outcome than the shared password you started with.

This is the same discipline as writing the recovery checklist before you need it — the migration is a rehearsal, and the parts you cannot rehearse in a weekend should not be cut over in a weekend. See Write the Recovery Checklist Before You Need It for the pattern of documenting the path before the incident forces you to improvise it.

Audit evidence that the shared credential is actually retired

Retiring the shared password is not the same as deleting it from the password manager. The credential may still exist in CI variables, in a Terraform state file, in a developer’s ~/.pgpass, or in a runbook. The verification step is to look for evidence of use, not evidence of deletion.

Three sources, in order of usefulness:

  1. PostgreSQL connection logs. If log_connections is on, every successful login is recorded with the role name. Query the logs for the old shared username over a two-week window after cutover. Any hit is a system you missed.
  2. Vault audit logs. Vault’s audit device records every request, including credential issuance. The OWASP auditing guidance applies here: you want to see who requested which role, when, and whether the request succeeded. If your audit device is not enabled before the migration, enable it first — a migration with no audit trail is a migration you cannot verify.
  3. Access review. A short written review, dated, listing each engineer, their Vault entity, their policies, and the PostgreSQL roles their Vault roles map to. This is the artifact that answers “who can reach production data” without a meeting.

NIST SP 800-53 Rev. 5 is the reference catalog if your organization needs a control mapping; the Access Control, Audit and Accountability, and Identification and Authentication families are the relevant ones. For a five-person team, the practical value is the vocabulary, not the full control set.

Onboarding and offboarding after the change

The measurable win of this migration is that offboarding becomes a single action. Before, removing someone meant finding every place the shared password was copied. After, it means disabling their Vault entity and, if you use an IdP-backed auth method, disabling their IdP account. Their dynamic credentials expire on their own; there is nothing to rotate.

Onboarding is similarly short: create the Vault entity, attach the policy that allows reading the appropriate database role, and confirm the engineer can issue a credential and connect. The PostgreSQL side needs no per-person role creation because the Vault role’s creation statement handles it. The one thing that does not disappear is the access review — a new engineer with oncall_write is a decision worth recording.

The honest tradeoff: dynamic credentials mean that a long-running interactive session will be interrupted when its TTL expires. Set the default TTL to something that matches how your team actually works — an hour is a reasonable starting point for interactive access, and the documentation example uses exactly that — and document the renewal command in the runbook. An engineer who does not know how to renew a lease during an incident will reach for the break-glass path, which defeats the design.

Questions that come up

Do we need Vault Enterprise? No, for dynamic PostgreSQL credentials. The rootless static-role workflow is Enterprise-only and does not support dynamic roles, so Community users configure the database secrets engine with a privileged connection user. That user becomes a credential to protect carefully.

What TTL should we use? The documentation example uses default_ttl="1h" and max_ttl="24h". For interactive human access, an hour is a reasonable default; for CI, match the deploy cadence; for break-glass, shorter than either. These are starting points, not findings.

Can we keep the shared password as a fallback? Only if it is rotated, stored separately from Vault, and its use is logged and reviewed. A fallback that is never rotated is the original problem with a new name.

What if Vault is down? The break-glass path above is the answer. Test it once before you need it, and record the test in the runbook.

How do we know the migration worked? The old shared username stops appearing in PostgreSQL connection logs, Vault audit logs show individual credential issuance, and the access review lists every engineer by name. Three artifacts, all checkable.

At 02:14 on a Saturday, a backup job fails. The error is not a disk failure or a network partition. It is an authentication error: the token used by the backup service account has expired, and no one on the on-call rotation knows which human owns the account, which system issued the token, or how to rotate it without breaking the restore path. The incident is not a breach. It is a mapping failure.

For lean technical teams — 2 to 15 engineers who are their own on-call rotation, running production on AWS, GCP, or bare metal with no dedicated SRE department — non-human credentials are the quiet majority of identities. They outnumber human users, they rarely appear in onboarding checklists, and they are often discovered only when they stop working. This article is about mapping them before that Saturday.

What counts as a non-human credential

A non-human credential is any identity used by a workload, job, or automation rather than a person. In practice, that includes:

  • AWS IAM roles attached to EC2 instances or ECS tasks, which supply temporary credentials that update automatically before expiry (AWS documentation, Use an IAM role to grant permissions to applications running on Amazon EC2 instances).
  • GCP service accounts, which are identified by email address and can authenticate as themselves or be impersonated by other principals (Google Cloud documentation, Service accounts overview).
  • PostgreSQL WAL archive commands that run under the ownership of the same user as the PostgreSQL server, and therefore depend on that user’s environment and credentials (PostgreSQL documentation, Continuous Archiving and Point-in-Time Recovery).
  • Restic repository passwords and access keys used by forget and prune operations, which lock the repository during pruning and require credentials that can delete data (restic documentation, Removing backup snapshots).
  • PagerDuty escalation keys, synthetic check API tokens, and monitoring webhook secrets.

The common thread is that each credential has a lifecycle: creation, use, rotation, and revocation. If no human is named as its owner, that lifecycle is unmanaged.

Why ownership mapping fails in lean teams

In a team of five engineers, the person who created a service account is often the only one who knows why it exists. When that person leaves, the account becomes an orphan. The failure mode is not dramatic. It is a slow accumulation of accounts that no one can safely delete because no one can prove they are unused.

NIST SP 800-53 Rev. 5 provides a control catalog that includes identification and authentication controls, but it does not prescribe a specific inventory format for non-human credentials. The control family structure implies that credentials should be managed, but the operational work of mapping them falls to the team. The CISA/NSA Enduring Security Framework guidance for software supply chain security notes that dependencies that are inadequately communicated or addressed may lead to vulnerabilities, and that poor supplier enterprise or development hygiene is a risk area. The same logic applies internally: an undocumented service account is an uncommunicated dependency.

The practical consequence is that token expiry becomes an incident rather than a maintenance task. AWS documentation states that temporary credentials from instance metadata update automatically before they expire, and that applications caching those credentials should refresh them every hour or at least 15 minutes before expiry. That automatic refresh works only if the application is running and the instance profile is intact. If the instance is stopped, the role is detached, or the application caches credentials beyond the refresh window, the next start can fail.

GCP service accounts present a different failure mode. Google Cloud documentation warns that service account keys are a security risk if not managed correctly, and recommends short-lived credentials or impersonation where possible. But keys still exist in many lean environments because they are easy to create and hard to inventory. A key created for a one-time migration can remain active for years if no owner is named.

A minimum viable inventory

An inventory does not need to be a database. A version-controlled YAML file or a spreadsheet with a defined schema is sufficient if it is updated as part of the change process. The fields that matter are the ones that answer the Saturday question: who do I call, what does this do, and how do I rotate it?

Recommended fields:

  • Credential identifier: the ARN, service account email, or secret name.
  • Human owner: a named engineer, not a team alias. This is the person who can authorize rotation or deletion.
  • Purpose: one sentence describing what workload uses it and why.
  • Creation date: when it was created, and by whom if known.
  • Rotation cadence: how often it should be rotated, or whether it is automatically rotated by the platform.
  • Last-used evidence: a link to a log query, a CloudTrail event, or a monitoring check that shows recent use. This is the field that prevents orphan accumulation.
  • Dependencies: what breaks if this credential is revoked. For a PostgreSQL WAL archive command, that includes the archive storage path and the restore procedure. For a restic repository, that includes the prune schedule and the lock behavior during pruning.
  • Break-glass procedure: how to rotate or revoke the credential outside normal change windows.

The last-used evidence field is the one most often omitted. Without it, deletion is a guess. With it, deletion is a decision.

Mapping the backup path specifically

Backup credentials deserve separate attention because they are both critical and often over-privileged. PostgreSQL continuous archiving requires a continuous sequence of archived WAL files that extends back at least as far as the start time of the base backup. The archive command runs under the ownership of the PostgreSQL server user, and PostgreSQL documentation advises that the archived data be protected from prying eyes, for example by archiving into a directory without group or world read access. That means the credential used by the archive command must have write access to the archive location, and the human owner of that credential must be known.

Restic adds another layer. The forget and prune commands require credentials that can delete data. The documentation notes that during a prune operation, the repository is locked and backups cannot be completed. It also advises running restic check after pruning to detect damage to internal data structures. If the credential that runs prune is owned by no one, a failed prune can leave the repository locked or the retention policy unenforced. The inventory should record which credential runs prune, when it runs, and who is responsible for verifying the result.

For a practical procedure on writing the recovery steps before an incident, see Write the Recovery Checklist Before You Need It. That article covers the checklist itself; this one covers the credential ownership that makes the checklist executable.

Offboarding and break-glass

When an engineer leaves, the offboarding checklist typically covers their human account, their email, and their laptop. It rarely covers the service accounts they created. The failure mode is not immediate. It appears weeks or months later when a token expires or a rotation is due.

The control that addresses this is a pre-offboarding review of the inventory filtered by owner. Before the last day, the departing engineer should transfer ownership of each credential to a named remaining engineer. If a credential has no remaining use case, it should be revoked as part of offboarding, not left for later.

Break-glass access for non-human credentials is a separate problem. AWS documentation describes the iam:PassRole permission, which controls which roles a user can pass to an EC2 instance. Restricting this permission is a way to prevent a user from launching an instance with more permissions than they have. In a lean team, the break-glass procedure should specify which human can assume which role, under what conditions, and how the action is logged. The inventory should record the break-glass path for each critical credential, not just the normal rotation path.

What evidence would show this works

The claim that an inventory prevents Saturday token-expiry incidents is not a research finding. It is an operational hypothesis. The evidence that would support it is specific and local:

  • A reduction in the number of authentication-failure alerts that occur outside business hours.
  • A decrease in the time to resolve credential-related incidents, measured from alert to rotation.
  • An increase in the proportion of credentials with a named human owner and a last-used evidence link.
  • A successful restore lottery — a scheduled test where a random backup is restored using only the documented procedure and the inventoried credentials.

These are not benchmarks. They are the metrics a team can collect from its own incident log and inventory. The point is to make the mapping visible before the token expires, not to prove a universal rule.

FAQ

How often should non-human credentials be rotated?
There is no single answer. AWS IAM roles for EC2 supply temporary credentials that update automatically, so rotation is handled by the platform. GCP service account keys do not rotate automatically and should be rotated on a schedule the team defines, or replaced with short-lived credentials where possible. The inventory should record the cadence and the owner responsible for enforcing it.

What is the minimum viable inventory for a team of five?
A version-controlled file with the fields listed above. The critical fields are human owner, purpose, last-used evidence, and break-glass procedure. If those four are present and current, the inventory is useful. If they are absent, it is a list of names.

How do we handle a service account that no one claims?
Treat it as a finding. Query the last-used evidence. If there is no evidence of use in a defined period, propose revocation. If there is evidence but no owner, assign an owner before the next rotation window. The goal is not to delete everything unclaimed; it is to ensure that every credential has a human who can answer for it.

Does this apply to bare-metal environments?
Yes. The credential may be an SSH key, a database password, or a restic repository password stored in a configuration file. The ownership question is the same: who rotates it, who revokes it, and who is called when it fails at 02:14 on a Saturday.

Sources

A controlled failover test can reveal what your DNS setup actually does. The first 60 seconds after stopping the endpoint are not a single event. They are a sequence of independent caches, resolvers, health checks, and client retry policies, each with its own timer. For a lean team running production without a dedicated SRE department, the goal of a DNS failover rehearsal is not to prove the failover works. It is to measure how long the failover takes, which clients recover first, and which clients keep talking to a dead address.

This article is a proposed observation plan, not a report of a test Gray Haven Lab performed: choose a low-traffic window, stop a test endpoint under an approved change plan, and observe the recovery curve for at least 60 seconds. We will cover what DNS failover means in practice, why TTL is not a guarantee, how health checks interact with authoritative DNS, and what to record so the next rehearsal is faster. The companion piece, Write the Recovery Checklist Before You Need It, covers the pre-work that makes this rehearsal safe.

What DNS failover actually is

DNS failover is the practice of changing the answer an authoritative nameserver returns when a health check marks an endpoint unhealthy. It is not a load balancer. It is not a connection drain. It is a control-plane change that propagates through a distributed cache hierarchy at a speed you do not control.

Adjacent concepts matter here: record TTL, negative caching, resolver prefetch, client connection pooling, and health-check interval plus failure threshold. A DNS failover rehearsal exercises all of them at once. That is why the 60-second window is useful. Treat it as an initial observation window, not a promised recovery time or a guarantee of limited impact.

For a small team that runs its own on-call rotation, DNS failover may be a practical cross-region or cross-provider mechanism. It is also the mechanism most likely to be assumed rather than measured. The rehearsal exists to replace the assumption with a number.

The 60-second timeline, second by second

The following timeline is an observation template, not measured traffic or a statement of any provider’s defaults. Record your actual health-check interval, failure threshold, record TTL and client retry behavior before the test. Use the time bands below to organize observations, not to predict when your provider will switch an answer.

T+0s: you kill the healthy endpoint

At T+0, the endpoint stops answering. Stopping a process may produce an explicit connection failure; dropping network traffic may instead leave clients waiting for a timeout. Record the behavior you actually observe. Record which one you did. A process kill and a path blackhole produce different health-check behavior.

T+0s to T+30s: health checks accumulate failures

Watch the health-check results rather than assuming an interval multiplied by a failure threshold predicts the transition. Record the first failed check, the time the provider marks the endpoint unhealthy, and the first observed answer containing the failover target. Different observers or clients may see those events at different times.

During this window, clients are still receiving the healthy address. Some of them are failing. Some of them are retrying. Your error rate is already rising before DNS has changed anything.

T+30s to T+45s: authoritative DNS changes

Once the health check marks the endpoint unhealthy, the authoritative nameserver begins returning the failover target. This is a control-plane event. It does not push to resolvers. It waits for them to ask again.

If your record TTL is 60 seconds, a resolver that cached the healthy answer at T-1s will keep serving it until T+59s. If your TTL is 300 seconds, that resolver keeps serving the dead address until T+299s. This is the second surprise: the authoritative change is not the client-visible change.

T+45s to T+60s: the first wave of clients moves

Clients with short-lived resolvers, no caching, or aggressive retry logic move first. Clients behind corporate resolvers, ISP resolvers, or long-lived connection pools move later. There is no universal percentage of requests that will move by T+60s. Record the share on the failover target at T+60s, T+120s, and T+300s; the measured curve is the number worth keeping.

If you use a global accelerator or anycast front door, the timeline compresses because the client is not waiting on recursive resolvers. If you use plain DNS with a 300-second TTL, the timeline stretches. Neither is wrong. Both need to be known.

Why TTL is not a guarantee

TTL is a hint, not a contract. Resolver behavior and client-side caches can make the visible change differ from the record’s nominal TTL. Negative caching for NXDOMAIN and NODATA responses is governed separately by the SOA minimum TTL, whose actual configured value should be checked. If your failover depends on a fast negative-cache expiry, check the SOA minimum before the rehearsal, not during it.

Client-side DNS caching adds another layer. Runtime DNS cache settings and long-lived connection pools can extend the time before a client uses a new address. Inspect the effective settings of each important client rather than assuming its behavior. A DNS change does not reach a process that already has an open connection. This is why a rehearsal should include at least one long-lived client, not just curl loops.

Primary sources worth reading before you set TTLs: the IETF RFC 1035 definition of TTL and the RFC 2308 treatment of negative caching. Both are short and both explain why your 30-second TTL may not behave like 30 seconds.

Health checks are part of the failover, not a precondition

A common rehearsal mistake is to treat the health check as a binary gate. In practice, the health check is a timer with its own failure modes. A health check that depends on a single TCP port will mark unhealthy when the port closes, even if the application is still serving on another port. A health check that depends on an HTTP 200 will mark unhealthy during a slow database query, even if the endpoint is otherwise fine.

Before the rehearsal, write down the health-check interval, the failure threshold, the success threshold, and the protocol. After the rehearsal, compare the observed unhealthy transition time to the calculated one. If they differ by more than one interval, the health check is doing something you did not expect.

For teams using PagerDuty escalation, the rehearsal should also confirm that the failover event does or does not page. A DNS failover that pages the on-call engineer at 03:00 is a different operational decision than one that does not. Decide which you want before the rehearsal, not after.

What to record during the 60 seconds

The rehearsal produces a dataset, not a pass/fail. Record at least these fields:

  • Kill method: process stop, port close, or network blackhole.
  • Health-check interval, failure threshold, and protocol.
  • Record TTL and SOA minimum TTL at the time of the rehearsal.
  • Time to authoritative change, measured from the provider’s API or console.
  • Time to first client request on the failover target, measured from access logs.
  • Percentage of request volume on the failover target at T+60s, T+120s, and T+300s.
  • Error rate on the killed endpoint during the window.
  • Whether the event paged, and at what severity.

As a hypothetical diagnostic example, suppose a small share of requests still reaches the stopped endpoint after the record TTL has passed. Compare those requests by client and connection age. A long-lived connection pool or runtime cache is one possible cause; verify it in logs before changing TTL or retry settings.

Tradeoffs: shorter TTLs versus resolver load

Shortening TTLs reduces the failover window but increases query volume against your authoritative nameservers. For a small team, that tradeoff is usually acceptable for the specific records that participate in failover, and not acceptable for every record in the zone. Scope the change.

Another tradeoff: aggressive health checks detect failure faster but generate more false positives during transient network blips. A 10-second interval with a 1-failure threshold will mark unhealthy on a single dropped packet. A 10-second interval with a 3-failure threshold tolerates two. Neither is universally correct. The rehearsal tells you which one matches your error budget.

If you use a managed failover product, read the provider’s documented behavior for health-check evaluation and failover timing. AWS documents Route 53 health-check behavior in its DNS failover documentation. Cloudflare documents its load-balancing health checks in its health check documentation. These are the primary sources for the timers you are about to measure.

Pre-mortem questions before the next rehearsal

A pre-mortem is cheaper than a postmortem. Before the next rehearsal, answer these:

  • What is the worst thing that happens if the failover target is also unhealthy?
  • Which clients are known to cache DNS beyond TTL?
  • Does the failover target have capacity for the full production load, or only a fraction?
  • Who is authorized to trigger the rehearsal, and who is authorized to abort it?
  • What is the rollback path if the failover target degrades during the window?

These questions belong in the recovery checklist, not in the incident channel. The recovery checklist article covers how to structure them so they are usable under pressure.

Access hygiene around the rehearsal

A DNS failover rehearsal touches production DNS. That means it touches credentials. Before the rehearsal, confirm that the engineer running it has the minimum permissions required, that the break-glass path is documented, and that the change is logged in the same place as other production changes. After the rehearsal, confirm that any temporary credentials or elevated roles are revoked.

For teams with onboarding and offboarding processes that are still manual, the rehearsal is a good moment to test whether a departed engineer’s credentials would still work. If they would, that is a finding. Record it as a finding, not as a failure of the rehearsal.

Monitoring and failure rehearsal

Synthetic checks are one way to observe the client-visible failover curve. A synthetic check from two or three external regions, polling at a documented interval, can give you a time series showing when the failover target started answering. Pair it with a growth-rate alert on the failover target’s request rate, so you notice if the failover target is absorbing traffic faster than expected.

If you use PagerDuty escalation policies, decide whether the rehearsal should page. If it should not, suppress the alert for the window and record the suppression. If it should, treat the page as a real page and follow the escalation path. Both are valid. The invalid option is not deciding.

FAQ

How long should a DNS failover rehearsal last?

Start with a 60-second timeline, but keep observing until the authoritative answer and the important client paths have actually changed. With a 300-second TTL or a slow health check, that may require ten minutes or more. Record the stopping rule before the test.

Should I lower TTL before the rehearsal?

Lower it at least one old-TTL period before the rehearsal, so resolvers have had a chance to pick up the new value. If you lower it at T-0, many resolvers will still be serving the old TTL.

What if the failover target cannot handle full load?

Then the rehearsal should measure partial failover, not full failover. Record the capacity limit and the resulting error rate. A failover that works at 40% load and fails at 100% is a known limitation, not a surprise.

Does a global accelerator remove the DNS failover window?

It compresses it. Anycast and global accelerators move traffic at the network layer, so clients do not wait on recursive resolvers. The health-check and control-plane timers still apply, but the resolver-cache portion of the window is largely removed.

How often should we rehearse?

Choose a cadence that matches your change rate and risk. If the failover path changes, rehearse after the change. If the on-call rotation changes, rehearse after the change. Budget the test and its rollback as a real production change.

What to do next

Pick one low-traffic service with a DNS failover path. Write down the health-check interval, failure threshold, record TTL, and SOA minimum. Get an approved change window and rollback plan. Stop the chosen endpoint, record the timeline for as long as the configured timers require, and compare observed numbers with the plan. Then update the recovery checklist with what you learned.

The follow-up topic worth writing next is the client-side half of this problem: how long-lived connection pools and runtime DNS caches can extend the failover window beyond the DNS record’s TTL. Measuring that behavior helps explain whether a real failover feels like a brief interruption or a longer outage.

Why the Fix-or-Workaround Decision Matters for Lean Teams

A recurring failure is any incident that returns with the same signature: the same alert, the same service, the same approximate time window, the same customer-visible symptom. For a team of two to fifteen engineers who own their own on-call rotation, every recurrence forces a decision that has less to do with engineering taste and more to do with where the next hour of reliability budget goes. The choice is rarely binary. It is a portfolio decision: fix the code, absorb the failure with an operational workaround, or do both in sequence.

This article is a decision framework, not a mandate. It assumes you run production on AWS, GCP, or bare metal without a dedicated SRE department. It assumes your monitoring is PagerDuty or a comparable escalation tool, your backups are restic or pgBackRest, and your runbooks live somewhere a tired engineer can find them at 03:00. The goal is to make the fix-or-workaround call explicit, repeatable, and reviewable, so the same failure does not quietly consume a quarter of your on-call capacity.

Define the Failure Before You Define the Fix

Most bad decisions start with a vague incident description. “The database was slow again” is not a failure definition. A useful definition includes the trigger, the blast radius, the detection path, and the recovery path. Write it in one paragraph and keep it in the incident record.

The four fields that make a recurrence comparable

  • Trigger: What changed or accumulated before the failure? A cron job at 02:15, a batch import, a certificate rotation, a traffic growth rate above 20% week over week.
  • Blast radius: Which services, regions, or customer segments degraded? One availability zone, one read replica, one tenant.
  • Detection: Which synthetic check, growth-rate alert, or PagerDuty escalation rule caught it? If a human reported it first, that is a detection gap worth naming.
  • Recovery: What actually restored service? A restart, a failover, a restore from pgBackRest, a manual WAL replay, a config rollback.

Once two or more incidents share these four fields, you have a recurrence. Until then, you have noise. The distinction matters because workarounds applied to noise create permanent complexity for no reliability gain.

The Decision Framework: Five Questions in Order

Ask these questions in sequence. Stop at the first one that gives a clear answer. The order is deliberate: it front-loads the cheapest and most reversible options.

1. Is the failure already contained by an existing control?

If a runbook step, a circuit breaker, or a scheduled maintenance window already prevents customer impact, the recurrence may be acceptable. A nightly vacuum that briefly raises replica lag, caught by a synthetic check and absorbed by read routing, is not a code bug. It is a known cost. Document it, set a review date, and move on. The trap is treating every alert as a defect. Some alerts are the system telling you it is working as designed.

2. Does the workaround have a bounded lifetime?

Operational workarounds are legitimate when they are temporary and dated. A manual step in the runbook that says “restart the worker pool after the 02:15 batch” is acceptable for one quarter if the batch job is being replaced. It is not acceptable indefinitely, because it transfers reliability risk from the code to the on-call engineer’s memory. If the workaround has no expiry date, it is not a workaround. It is technical debt with a pager attached.

3. What is the cost of the code fix relative to the recurrence rate?

Estimate both sides in the same unit: engineer-hours per quarter. A failure that costs 45 minutes of on-call time twice a month is roughly 18 hours per quarter, plus context-switching and postmortem overhead. A code fix that takes 30 hours to design, implement, and verify pays back in under two quarters. A fix that takes 120 hours and touches a shared library used by six services may not pay back within the year, especially if the failure is cosmetic or self-healing.

Use your incident records, not intuition. If you do not have incident records, start with a simple table in your runbook repository: date, duration, detection source, recovery action. Three months of that table will change how you prioritize.

4. Does the fix reduce a class of failures or just one instance?

A code fix that eliminates one alert but leaves the underlying pattern intact is often worse than a workaround, because it creates false confidence. A fix that adds idempotency to a job runner, or that makes a retry policy respect backoff, may close a whole class of recurrences. Prefer fixes that generalize. When you cannot generalize, prefer workarounds that are visible and dated.

5. Can the workaround be tested like code?

If the workaround is a runbook step, it should be rehearsed. If it is a script, it should live in version control and run in a staging environment. If it is a manual database intervention, it should be paired with a restore drill so the team knows the recovery path is real. A workaround that has never been rehearsed is not a control. It is a hope. The Recovery Checklist Before You Need It is a useful template for turning ad hoc recovery steps into rehearsed procedures.

When a Workaround Is the Right Answer

Workarounds are not failures of engineering discipline. They are appropriate when the failure is rare, the impact is contained, the fix is expensive, and the workaround is observable. Three patterns justify a workaround:

  • Third-party dependency behavior: An upstream API returns intermittent 503s during its maintenance window. You cannot fix their code. You can add a retry with jitter and a synthetic check that distinguishes their maintenance from your outage.
  • Infrastructure-level noise: A cloud provider’s instance retirement notices cause a brief spike in your PagerDuty queue. The workaround is a suppression rule scoped to that event type, reviewed quarterly.
  • Low-frequency, high-cost fixes: A failure occurs once every 18 months and the fix requires a schema migration across a multi-terabyte PostgreSQL cluster. The workaround is a documented manual failover with a pgBackRest restore path, rehearsed twice a year.

In each case, the workaround has a named owner, a review date, and a detection signal. Without those three, it drifts into folklore.

When a Code Fix Is the Right Answer

Code fixes earn their cost when the failure is frequent, the workaround is fragile, or the fix closes a class. Signals that point toward a fix:

  • The same alert fires more than once per month and each occurrence consumes more than 30 minutes of on-call time.
  • The workaround requires a specific engineer’s knowledge, and that engineer is approaching a vacation or role change.
  • The failure has a customer-visible symptom, even if brief, and the workaround depends on a human noticing it.
  • The fix is localized: one service, one module, one configuration path, with a clear test that reproduces the failure.

When you choose a fix, write the decision record. Name the failure signature, the recurrence rate, the estimated fix cost, and the expected reduction. This is the same discipline as a blameless postmortem, applied before the work rather than after. A pre-mortem that asks “what would make this fix fail to reduce recurrences?” is cheaper than discovering it six weeks later.

Sequencing: Workaround First, Fix Second

For lean teams, the most common correct answer is both, in order. Ship the workaround this week to stop the bleeding. Schedule the fix for the next planning cycle. The workaround buys time; the fix buys capacity. The failure mode to avoid is shipping the workaround and never scheduling the fix, because the workaround made the pain invisible.

Make the sequence explicit in your runbook: workaround owner, fix owner, review date. If the fix slips, the review date forces a conversation rather than a silent deferral.

Access Hygiene and Break-Glass in the Decision

Some recurrences are not code or operations problems. They are access problems. A failure that requires a break-glass credential to resolve is a signal that your normal access path is insufficient. If the same break-glass account is used twice in a quarter, the recurrence is in your access model, not your application. Onboarding and offboarding hygiene, scoped roles, and time-bound break-glass credentials are part of the fix-or-workaround decision because they determine who can act at 03:00 and how much damage a mistake can cause.

Document break-glass use in the incident record. If the pattern repeats, treat it as a design input, not an operational quirk.

Monitoring and Failure Rehearsal as Decision Inputs

You cannot decide whether a recurrence deserves a fix if you cannot measure it. Synthetic checks, growth-rate alerts, and PagerDuty escalation rules are the instruments that turn a vague pattern into a comparable record. A growth-rate alert that fires when disk usage increases more than 15% in 24 hours is more useful than a static threshold, because it catches the trend before the outage. A synthetic check that exercises the restore path is more useful than a backup success notification, because it tests recovery rather than storage.

Rehearse the failure. If the workaround is a manual failover, run it in staging. If the fix is a retry policy, inject the failure and confirm the retry. Rehearsal converts assumptions into evidence, and evidence is what makes the fix-or-workaround call defensible.

FAQ

How many recurrences justify a code fix?

There is no universal number. A practical threshold for a lean team is two occurrences in a quarter with more than 30 minutes of on-call time each, or any occurrence with customer-visible impact and a fragile workaround. The threshold should be written down and reviewed, not rediscovered each time.

Can a workaround ever be permanent?

Yes, if it is bounded, observable, and cheaper than the fix over a multi-year horizon. A permanent workaround should have an owner, a review date, and a detection signal. If it lacks those, it is not permanent. It is unmanaged.

What if the fix is in a third-party dependency?

Then the decision is about mitigation, not repair. Add retries with backoff, synthetic checks that distinguish their maintenance from your outage, and a documented fallback. Track their incident history and review your mitigation quarterly. You cannot fix their code, but you can reduce your exposure to it.

How do we avoid workaround drift?

Put every workaround in the runbook with a date and an owner. Review the runbook quarterly. If a workaround has no review date, it is drift. The same discipline that keeps your backup and recovery drills current keeps your workarounds honest.

What to Do Next

Pick one recurring failure from the last quarter. Write its four fields: trigger, blast radius, detection, recovery. Ask the five questions in order. Record the decision, the owner, and the review date. Then rehearse the chosen path, whether it is a code fix or a workaround, so the next occurrence is a known procedure rather than a fresh incident. The decision framework is small. The discipline of applying it consistently is what keeps a lean team’s on-call rotation sustainable.

Consider a fictional postmortem titled “INC-2024-017: API latency degradation.” Six months later, a new teammate searching the incident index might still be unable to tell from that title what failed or which fix mattered. This worked example shows how a timeline, causal explanation, and precise title can make an incident record useful to someone who was not in the room. The team, outage, documents, times, and metrics below are invented to demonstrate the method; they are not Gray Haven Lab’s observed history.

A fictional incident to work through

Imagine a four-person team running a PostgreSQL primary with a streaming replica. In this invented scenario, a Tuesday 14:10 UTC page reports elevated 5xx rates on a checkout API; service returns to baseline around 14:50. The times and impact are illustrative, not measurements from a real outage. They give the postmortem exercise a concrete sequence to test.

Suppose the exercise provides chat messages, application logs, a dashboard capture, and an alarm history. The task is to determine what each artifact can establish, where the first draft overstates a conclusion, and what a later reader would need to verify it.

Chronology is not a narrative

A first draft can become a flat timeline: timestamp, event, timestamp, event. It may record the sequence without explaining which events changed the outcome. Accuracy matters, but a useful postmortem also identifies the failure mechanism and the evidence for it.

A timeline entry such as “14:12 — on-call acknowledged page” records an event. A stronger entry would say “14:12 — page acknowledged; responders checked the load balancer first because the alert did not name the checkout API”, if the chat and logs support that account. The second version tells a future reader why those minutes mattered. The writing analogy is a way to test the explanation, not evidence that the incident occurred.

For this fictional case, organize the evidence into three questions rather than forcing every timestamp into an act structure:

  • Before the alert. What was normal, and which condition made the failure possible? The invented records show a connection pool approaching its limit before a deployment increased query traffic. A real postmortem would need metrics to establish the timing.
  • The turning point. In the worked example, pool exhaustion queues health checks, targets are marked unhealthy, and traffic shifts to fewer app nodes before the page fires. The alert is an observation of the failure, not necessarily its first cause.
  • Resolution and follow-up. Record the action taken, the measurements that showed recovery, and the permanent change proposed. Do not treat “error rate returned to baseline” as a substitute for evidence that the cause is understood.

This analysis takes more time than pasting a timeline. Spend that time where the incident reveals a new failure mode or meaningful customer impact. For a repeat of a known issue, a shorter record may be enough if it still links the evidence, decision, and follow-up.

Reconstructing the timeline from Slack, not memory

People in the same incident can remember its order differently. Reconstruct the sequence from available artifacts and record their limits:

  1. Preserve the incident channel within the organization’s retention rules. Keep message timestamps and relevant context, subject to access and privacy controls. Chat is one clock, not the only authoritative record.
  2. Compare chat with application logs and metrics. Chat shows what responders believed; logs and metrics show what systems recorded. Neither is infallible. In this fictional checkout case, the chat focuses on a recent deploy while the sample metrics place pool saturation earlier. A real report would show the timestamps and queries behind that conclusion.
  3. Mark belief changes explicitly. Annotate each working hypothesis and the evidence that supported or challenged it. For example: “14:19 — hypothesis: bad deploy; counterevidence: pool metric already saturated at 13:58.” In a real report, attach the underlying metric and note its clock source.
  4. Separate the incident timeline from the communication timeline. Record when status updates went out, what they said, and which observations supported them. Parallel columns make any gap visible without inventing motives.

One caution: a postmortem can inherit the tone of a hurried chat. Keep supported claims about systems; remove speculation about people. “The dashboard was green, so responders initially checked another service” can be a finding if the artifacts show it. “Dana was slow” is a judgment that does not identify a mechanism.

Naming the failure mode like a chapter, not a ticket

Here’s where the craft matters most. Compare two titles for the same incident:

  • “INC-2024-017: API latency degradation”
  • “The Green Dashboard That Hid a Saturated Connection Pool for Two Weeks”

The first is a filing label. The second is a compressed argument: it names the misleading signal (the dashboard), the real failure (pool saturation), and the duration of the hidden condition (two weeks). Someone scanning the postmortem index six months later can tell from the second title that this is the document to read before they trust their own replica dashboards. The first title tells them nothing except that an incident existed.

A useful title rule is to name the failure mode, mechanism, and, when supported, the surprise. “Replica promoted but lost the last 40 seconds of writes” and “Health check passed while the pool behind it was exhausted” are hypothetical examples of that form. “Database issues on Tuesday” is too vague to guide a later search.

Incident titles can benefit from the same revision question as any short title: what central tension should a reader understand before opening the document? Draft a few options from the verified failure mechanism, then test whether each one names a cause instead of a symptom.

Before publishing a real postmortem, ask someone who was not on the incident whether the title alone indicates the category of failure and the first system they would check. If it does not, revise the title after the timeline and root-cause evidence agree. The title helps readers locate the report; its precision should come from the investigation.

What the worked example suggests changing

The fictional record suggests three revisions. First, write the title after reconstructing the timeline, so the first hypothesis does not become the headline. Second, include enough pre-alert metrics to show whether an enabling condition existed before the page fired. Third, keep the narrative concise: split independent failure modes into separate findings instead of letting an elaborate story hide the evidence.

A postmortem serves someone who was not present: a new hire, a future responder, or an engineer deciding whether to repeat an architectural choice. Chronology is necessary but not sufficient. A clear causal account, linked to artifacts and titled for the verified failure mode, gives that reader a way to test the lesson.

Every team that runs production long enough collects failures that keep coming back: a nightly restic job that stalls on a locked index, a PostgreSQL connection pool that saturates under one specific report, a pgBackRest archive gap that shows up after weekend batch work. Recurring-failure triage — the decision to spend engineering hours on a code fix or to keep the incident behind an operational workaround such as a runbook step, a cron script, or a widened alert threshold — is a routine call that quietly shapes reliability on a small team. The adjacent concepts matter as much as the verdict itself: error budgets, MTTR trends, workaround debt, blameless postmortems, and decision records. For a team of two to fifteen engineers with no dedicated SRE function, engineering hours are the scarcest resource in the system, and every workaround is a standing tax on on-call attention. This is the method we use. Thresholds included.

What counts as a recurring failure

Answer first: three occurrences of the same root cause inside a rolling ninety days is the working definition of recurring. Symptom clustering is not evidence. Three pages tagged latency can be three different diseases with three different correct responses, and treating them as one recurring failure is how a team ends up fixing the wrong thing twice.

Symptoms lie. A six-person fintech we will call Corvid kept paging on 5xx spike; incident review showed the three pages in one month were a CDN misconfiguration, a connection pool exhausted by a reporting query, and a slow migration holding locks. Three causes, three decisions — one vendor ticket, one code fix, one scheduling change. Tagging by root cause, not by alert name, is what makes the count mean anything.

The comparison step is a blameless postmortem exercise — the discipline Google’s SRE practice formalized in the SRE Workbook chapter on postmortem culture. On a small team the postmortem can be thirty minutes and a shared doc. The point is to compare causes, not to produce ceremony.

Below three occurrences, document and wait. Two is a coincidence worth a line in the incident record. One is just an incident. At three, the decision goes on the ledger — fix, workaround, or monitor — whether or not the answer feels obvious yet.

Tangled cables and status lights on network hardware in a server rack
Recurrence is a property of root cause, not of the alert name.

The decision frame: carrying cost versus fix cost

Promote a workaround to a code fix when the workaround’s carrying cost compounds faster than the fix’s one-time cost. Everything else in this method is detail.

Carrying cost has four line items. Frequency times minutes per occurrence times people involved. Pager load, because every page spends on-call attention even when the fix is easy. Execution-error risk, because manual steps have a failure rate and it is highest at 3 a.m. And drift, because workarounds rot as the system changes around them. Fix cost has three: engineer-hours, regression risk, and deploy risk on whatever path the change lands.

Concrete numbers make the tradeoff visible. A workaround firing every eleven days with a forty-minute MTTR and two people involved costs roughly 3.5 hours a month before you count the pager. A fix estimated at twelve engineer-hours pays for itself in under four months if it removes the failure outright, and in eight if it only halves the frequency. If the same fix has to land on the payment authorization path, the regression risk may dominate both numbers — which is why the frame has two sides.

When the code fix wins

The failure touches a recovery path

Any known defect in backup, restore, or WAL archiving gets a fix, not a workaround. Recovery paths get exercised at the exact moment the system is already down; a workaround there is a bet placed during your worst hour. restic check reporting index inconsistencies, a pgBackRest archive stall on PostgreSQL 16, a WAL gap discovered mid-restore — these are promotion cases with no counting period. PostgreSQL’s durability model depends on write-ahead log replay, and the WAL chapter of the PostgreSQL documentation is blunt about the ordering guarantees you are relying on. A workaround that leaves those guarantees unverified is not a workaround. It is an untested restore.

Detection is cheap. A monthly restore lottery — restore a randomly selected backup to a scratch host and time the result — catches workarounds that quietly rotted. If a workaround touches restore steps at all, those steps belong on a written recovery checklist, not in tribal memory; the approach is covered in Write the Recovery Checklist Before You Need It.

The workaround depends on someone’s memory at 3 a.m.

If the workaround’s steps live in one person’s head, that is a fix signal, not a runbook. The test is escalation behavior: when the PagerDuty escalation for this incident converges on a specific engineer — only Dana knows the restart order — the workaround is a single point of failure wearing a runbook costume. Promote it to code, or at minimum to a tested runbook step behind a synthetic check. Steps that require judgment under fatigue have an error rate. We have watched a correct six-step procedure get executed wrong twice in one quarter because step four depended on reading a stale Grafana dashboard.

Frequency and impact cross a threshold you wrote down in advance

Pick the thresholds in advance and write them down; do not negotiate them per incident. The set we use in these field notes: same root cause three times in ninety days with MTTR over thirty minutes — fix. Any occurrence with data-loss exposure — fix regardless of count. Frequency rising two months in a row — fix review, even if MTTR is short. Below those lines, a documented workaround is a defensible answer.

Code and terminal output open on a laptop screen
The promotion test: carrying cost, memory dependency, blast radius.

When the workaround wins

You do not own the failing layer

If the defect sits inside a managed service, the workaround is often the only lever you actually hold. Cloud SQL failover quirks, intermittent S3 slow-down responses during heavy list operations, a GCP load balancer behavior that only appears under one traffic shape — file the vendor case, keep the runbook step, and record the case ID in the ledger. The tradeoff: waiting on a vendor is a decision, not a default, so it gets a revisit date like any other workaround.

The fix’s blast radius exceeds the failure’s cost

Compare the fix’s regression risk against the failure’s measured cost, not against its annoyance. A four-person payments team we will call Heron kept a manual failover runbook for a flaky connection pool rather than touching the transaction path mid-quarter; the fix shipped later, in a planned window, with a rollback plan and a canary. The classic version of this branch is the PostgreSQL major-version upgrade that would remove the failure class entirely. The workaround is legitimate there — as a bridge with an end date, not as a residence.

The failure is bounded and the error budget is not

A bounded failure that stays inside its error budget is a candidate for a documented workaround. Bounded means capped blast radius, self-limiting duration, and no data exposure. If the monthly burn-rate alerts in Grafana still finish the month in the black, the workaround is defensible; the moment burn alerts fire two months running, the promotion review starts. This is the one branch where doing nothing expensive is often the correct answer — provided the ledger entry exists.

Write the decision down: the workaround ledger

Every accepted workaround gets a one-page decision record; undocumented workarounds are just failures you have agreed to forget. The format borrows from architecture decision records — context, decision, consequences — trimmed to what an on-call engineer will actually read:

  • Incident IDs (PagerDuty references) that triggered the decision
  • Root cause in one sentence
  • Workaround steps, in runbook form
  • Carrying-cost estimate: frequency, MTTR, people involved
  • Promotion trigger — the threshold that would force a fix
  • Revisit date and owner

The ledger review is a fixed thirty-minute quarterly slot with three questions: which entries fired since last review, which missed their revisit date, and which thresholds should change. Entries that miss their revisit date get promoted or retired. The ledger is allowed to shrink, and most quarters it should.

Small team talking through an incident review around a conference table
The quarterly ledger review: promote, retire, or adjust the thresholds.

A worked example from the field

A nine-person B2B SaaS team on AWS — anonymized here as Kiln — ran restic 0.16.4 against a 4 TB repository, and the nightly prune failed roughly every eleven days with an index inconsistency (pack file cannot be found) that on-call resolved with restic rebuild-index and a retry. MTTR was about forty minutes, two people were usually involved, and the window produced three pages in ninety days.

The ledger entry made the arithmetic plain: about 3.5 hours a month of on-call attention, on a recovery path, crossing the three-in-ninety-days line. The fix — moving prune to a weekly systemd timer with a restic check pass beforehand, plus explicit lock cleanup — was estimated at ten engineer-hours and paid for itself inside a quarter. It shipped; occurrences went to zero and stayed there for two quarters. The follow-up mattered as much as the fix: the next restore lottery caught a stale copy of the old rebuild-index runbook step still linked from the wiki, and deleting it removed exactly the memory-dependent failure this method is designed to catch.

The counterfactual is worth stating. Had the same failure lived on the vendor’s side of the fence — say, inside a managed database’s snapshot scheduler — the same ledger entry would have justified keeping the workaround, with a case number and a revisit date instead of a sprint ticket. The method does not bias toward fixes. It biases toward deciding once, with numbers, on paper.

Frequently asked questions

How many recurrences justify a code fix?

Three occurrences of the same root cause in a rolling ninety days is a reasonable promotion trigger, provided MTTR exceeds about thirty minutes or the workaround needs more than one person. Any recurrence that touches backup, restore, or WAL archiving skips the count entirely and goes straight to a fix, because a workaround on a recovery path is a bet placed during your worst hour.

Is a runbook step enough, or is that just a workaround with better documentation?

A runbook step is the floor, not the finish line. It removes the memory dependency but none of the carrying cost — the pages, the minutes, the execution risk. If the step still fires monthly, it remains a promotion candidate. Runbooks are where workarounds wait; they are not where workarounds retire.

How do you track workaround debt across quarters?

One ledger file, one page per workaround: incident IDs, root cause, steps, carrying-cost estimate, promotion trigger, revisit date, owner. Review in a fixed thirty-minute quarterly slot. Any entry that misses its revisit date gets promoted or retired. The ledger should shrink most quarters; if it only grows, your promotion thresholds are too patient.

When is a monitoring change enough — no fix and no workaround?

When the failure is genuinely rare, self-limiting, and cheap to confirm, a synthetic check or an adjusted alert threshold can be the entire response. The test: if the check fired next month, would anyone change behavior? If no, you are collecting noise. If yes, you have a workaround by another name, and it belongs on the ledger.

Who makes the promotion call on a small team?

The engineer who carries the pager for that service proposes; the weekly ops slot confirms, with one other engineer as a sanity check. That avoids both failure modes of small-team decisions — the solo hero fix at 2 a.m., and the silent endurance of a workaround nobody officially accepted.

Where this goes next

The promotion decision and the recovery checklist are the same muscle: both are about writing down, in advance, what you will do when the system is at its worst. The companion piece is Write the Recovery Checklist Before You Need It, which covers the restore-path side of this method. And if you are holding a workaround you cannot decide about, send the ledger entry — anonymized recurring failures are the raw material of this column.