Every lean team has a version of this conversation. A new engineer joins on-call, opens the primary database, and asks why the service runs PostgreSQL instead of the managed alternative the rest of the stack uses. The answer exists. It lives in the head of the person who made the call, in a Slack thread from two years ago, and in a half-finished design doc that was never merged. When that person leaves, the answer leaves with them.
A one-page decision record is the smallest artifact that prevents this. It is not a design document, not a runbook, and not a postmortem. It is a dated, immutable note that captures one choice, the constraints that made it reasonable, and the conditions under which it should be revisited. For a team of two to fifteen engineers who are their own on-call rotation, it is the cheapest form of institutional memory available.
What a decision record is, and what it is not
The concept is not new. The Architectural Decision Record community defines an architectural decision as “a justified design choice that addresses a functional or non-functional requirement that is architecturally significant,” and an ADR as a record that “captures a single AD and its rationale.” The same source notes that the practice extends beyond architecture to “any decision record.” Michael Nygard’s 2011 blog post popularized the format, and the community maintains a comparison of seven templates.
What matters for a small team is the constraint on size. A decision record is not a design doc. It does not enumerate every alternative in depth, does not contain implementation steps, and does not get updated as the system evolves. It is a snapshot of reasoning at a point in time. If you find yourself writing more than a page, you are writing a design doc, and it will not get read during an incident.
It is also not a runbook. A runbook tells you how to restore the database. A decision record tells you why the database is PostgreSQL, what you gave up by choosing it, and what would have to change for the choice to be wrong. The two documents serve different readers at different moments. The runbook is read at 3 a.m. during a restore. The decision record is read during onboarding, during a migration debate, and during the postmortem that asks whether the original choice contributed to the incident.
The fields that earn their place on one page
A useful decision record for a lean team has seven fields. Each one exists because its absence causes a specific failure mode.
Title and date. A short imperative title (“Use PostgreSQL 16 as primary datastore for billing service”) and the date the decision was made. The date matters because it anchors the decision to a version of the system and a set of constraints that may no longer hold. PostgreSQL 18.6 is the current documented release as of this writing; a decision made against PostgreSQL 12 should be read with that gap in mind.
Status. One of: proposed, accepted, superseded by [link], or deprecated. This is the field that keeps the record honest. A decision record that is never superseded becomes folklore. A decision record with a status field can be retired cleanly.
Context. The operational reality at the time. Team size, on-call rotation, budget ceiling, compliance obligations, existing skills, and the scale the service was expected to handle. This is the field most often skipped and most often needed. “We chose PostgreSQL because it was free” is not context. “We chose PostgreSQL because the team had two engineers with production PostgreSQL experience, the managed alternative cost $X per month at our projected write volume, and we needed row-level security for a compliance requirement that the alternative did not support at the time” is context.
Decision. One or two sentences stating what was chosen. No hedging.
Alternatives considered. A short list with the reason each was rejected. This is where the record earns its keep during a future migration debate. If the team later considers moving to a managed service, the record shows whether the original objection still applies. If the objection was cost at a specific volume, and volume has changed, the record tells you the decision is ripe for revisit.
Consequences. What the team accepted by making this choice. Operational burden, backup and recovery responsibility, upgrade cadence, the need for a WAL archiving strategy, the need for a tested restore procedure. This is the field that connects the decision record to the recovery checklist. If the decision to run PostgreSQL yourself means you own point-in-time recovery, the decision record should say so, and the recovery checklist should reflect it.
Revisit triggers. The conditions under which this decision should be reopened. Specific, observable conditions. “If write volume exceeds X per second,” “if the team drops below two engineers with PostgreSQL experience,” “if the managed alternative adds row-level security at a cost below Y.” Without this field, the decision record is a historical document. With it, the record is a live input to planning.
Why PostgreSQL specifically makes this exercise worthwhile
PostgreSQL is a reasonable default for many small teams, and that is exactly why the reasoning behind choosing it is easy to lose. The choice feels obvious in retrospect, so nobody writes it down. Then a new engineer arrives from a shop that ran MySQL or a managed document store, and the obviousness is no longer shared.
The operational surface of PostgreSQL is large enough that the decision has real consequences. The PostgreSQL documentation devotes chapters to backup and restore, high availability and replication, reliability and the write-ahead log, and monitoring database activity. Each of those chapters represents work that a lean team either does or explicitly decides not to do. A decision record that says “we chose PostgreSQL” without saying which of those responsibilities the team accepted is incomplete.
Configuration is another area where the original reasoning matters. Tools like postgresqlco.nf exist because PostgreSQL’s configuration parameters are numerous and their interactions are non-obvious. A decision record does not need to capture every setting, but it should capture the constraint that drove the initial configuration posture. If the team chose to run with default settings because the workload was small and predictable, that is a decision. If the team tuned for a specific write pattern, that is a decision. Either way, the next person needs to know which one it was.
Where the record lives, and how it is found
A decision record that cannot be found during an incident is not a decision record. It is a file.
The storage location should be the same place engineers already look for project context. For most teams, that is the repository. A docs/decisions/ directory with files named YYYY-MM-DD-short-title.md sorts chronologically and is discoverable with a directory listing. The GitHub documentation on READMEs notes that a README is often the first item a visitor sees and that GitHub surfaces a README from the .github, root, or docs directory. A short pointer in the repository README to the decisions directory costs one line and saves a search.
Naming matters more than format. A file named database-choice.md is ambiguous after three database-related decisions. A file named 2024-03-11-postgres-primary-datastore.md is self-describing. The date prefix also makes supersession visible: a newer file with a related title is easy to spot.
Linking matters as much as naming. The decision record should link to the runbook that covers backup and restore, and the runbook should link back to the decision record. The postmortem template should include a field for “relevant decision records.” The onboarding checklist should include a pass through the decisions directory. Each of these links is cheap to add and expensive to reconstruct later.
How decision records interact with postmortems and runbooks
The Google SRE book chapter on postmortem culture describes a postmortem as “a written record of an incident, its impact, the actions taken to mitigate or resolve it, the root cause(s), and the follow-up actions to prevent the incident from recurring.” It also notes that postmortems are expected after any significant undesirable event, and that common triggers include user-visible downtime, data loss, on-call intervention, and monitoring failure.
A decision record is not a postmortem, but the two documents should reference each other. When an incident traces back to a database configuration choice, the postmortem’s root cause analysis should cite the decision record that established the configuration. When the postmortem produces an action item to change the configuration, that action item should produce a new decision record or a supersession of the old one.
The SRE book also emphasizes that blameless postmortems focus on “identifying the contributing causes of the incident without indicting any individual or team.” A decision record supports this by making the original reasoning visible. If a decision that seemed sound at the time contributed to an incident, the record shows what information was available when the decision was made. That is the difference between “who chose this” and “what did we know when we chose this.”
Runbooks and decision records have a simpler relationship. The runbook is procedural. The decision record is contextual. A runbook that says “run pgBackRest restore” is more useful when the reader understands why pgBackRest was chosen over the built-in pg_basebackup and WAL archiving. The decision record provides that context without cluttering the runbook.
A minimal review process
The failure mode of decision records in small teams is not that they are written badly. It is that they are written once and never revisited, so they drift out of alignment with the system. The fix is a review cadence that is small enough to survive a busy quarter.
One approach: review the decisions directory once per quarter, during an existing team meeting. For each record, ask one question: is the status still accurate? If yes, leave it. If no, either supersede it with a new record or mark it deprecated with a one-line note explaining why. This is a fifteen-minute exercise for a team with a dozen records.
A second approach: tie review to events rather than calendar. When a new engineer joins, they read the decisions directory as part of onboarding and flag any record that does not match what they observe. When a postmortem produces an action item that changes a previously decided approach, the postmortem owner writes the superseding record. When a revisit trigger fires, the person who notices writes the new record.
Both approaches work. The important thing is that supersession is a normal, low-ceremony act. A superseded record is not a mistake. It is a record of a decision that was correct under the constraints that existed when it was made.
What this looks like in practice
A team of six engineers runs a billing service on PostgreSQL 16, self-managed on EC2. The original decision was made eighteen months ago by an engineer who has since left. The decision record, written at the time, says the team chose self-managed PostgreSQL over Amazon RDS because the projected write volume at the time would have pushed RDS into a cost tier the budget could not absorb, and because the team had two engineers with production PostgreSQL experience.
The consequences section notes that the team accepted responsibility for WAL archiving, point-in-time recovery, and minor version upgrades. The revisit triggers include “if the team drops below two engineers with PostgreSQL experience” and “if RDS pricing for the required instance class falls below the current EC2 cost plus 20 percent.”
Eighteen months later, one of the two experienced engineers has moved to a different team. The revisit trigger has fired. The current team reads the decision record, sees that the original cost objection may no longer apply, and opens a new decision record to evaluate the migration. The old record is marked superseded. The new record cites the old one. The reasoning is preserved across the transition.
None of this required a design doc. It required one page, written at the time of the decision, by the person who made it.
Frequently asked questions
How is a decision record different from a design doc? A design doc describes how a system will be built. A decision record describes why a specific choice was made and what was traded away. Design docs are often long and are written before implementation. Decision records are short and are written at the moment of decision. A design doc may contain several decision records, or none.
What if the decision was made years ago and nobody remembers the reasoning? Write the record now, with the date of the original decision if known, and note in the context section that the reasoning is reconstructed. A reconstructed record is better than no record. It also creates a natural moment to ask whether the decision still holds.
Should every decision get a record? No. The ADR community frames the threshold as decisions that address “architecturally significant” requirements, meaning requirements with a measurable effect on architecture and quality. For a lean team, a practical filter is: if a new engineer would need to know this to avoid making a mistake, write it down. Database choice, backup strategy, authentication approach, and deployment topology usually qualify. Library version bumps usually do not.
How do decision records relate to the recovery checklist? The decision record explains why the recovery approach was chosen. The recovery checklist is the procedure. If the decision record says the team accepted responsibility for point-in-time recovery, the recovery checklist should include a tested PITR procedure. The two documents should link to each other.
What happens when a decision record is wrong? Mark it superseded, write a new record that explains what changed, and link the two. The old record stays in place. Its value is in showing what was known at the time, not in being correct forever.
Does this work for decisions other than database choice? Yes. The format is general. The adr.github.io project notes that ADR usage “can be extended to design and other decisions.” The constraint is the same: one page, dated, with context, consequences, and revisit triggers.