Writing a Postmortem When the Root Cause Is You: A Blameless Template for the Deploy You Made Yourself

On a team of two to fifteen engineers, the person who pushed the change, the person who answered the page, and the person who writes the postmortem are usually the same person. That is not a moral problem. It is a documentation problem with a known failure mode: the postmortem either becomes a confession, or it does not get written at all.

The sources are clear about what a postmortem is for. Google’s SRE book defines it as “a written record of an incident, its impact, the actions taken to mitigate or resolve it, the root cause(s), and the follow-up actions to prevent the incident from recurring.” It also states that “writing a postmortem is not punishment—it is a learning opportunity for the entire company.” Those two sentences sit in tension when you are the author of the change. The format names the actions that led to the incident; the culture forbids indicting the person who took them.

That tension is not a reason to soften the record. It is a reason to be more precise about where the causal chain actually runs.

Blameless does not mean actionless

The SRE book is explicit: “For a postmortem to be truly blameless, it must focus on identifying the contributing causes of the incident without indicting any individual or team for bad or inappropriate behavior. A blamelessly written postmortem assumes that everyone involved in an incident had good intentions and did the right thing with the information they had.”

Read that carefully. Blameless does not mean the actions disappear. The same chapter acknowledges that “blameless postmortems can be challenging to write, because the postmortem format clearly identifies the actions that led to the incident.” The actions stay in the document. What changes is the question you ask about them.

The question is not “why did I do that.” The question is “what did the system show me at that moment, and what would have made the correct action the default.”

This matters more for a self-authored postmortem, not less. The SRE workbook’s case study of a bad postmortem is blunt about the cost of individual highlighting: “It may seem like a good idea to highlight individuals in a postmortem. Instead, this practice leads team members to become risk-averse because they’re afraid of being publicly shamed. They may be motivated to cover up facts critical to understanding and preventing recurrence.” On a small team, the person most likely to cover up a fact is the person who wrote the deploy. If the postmortem reads as a confession, the next one does not get written.

A three-column causal chain

The practical move is to write the causal chain in three columns before you write prose. This is a working method, not a source-prescribed template, but it follows the sources’ emphasis on contributing causes over individual fault.

Column one: what the operator did. State it plainly. “Ran the migration command a second time after the first attempt returned a timeout.” No adjectives. No “carelessly” or “stupidly.” The workbook’s bad-postmortem example flags “careless ignorance” as superfluous language that distracts from the key message.

Column two: what the system showed at that moment. This is the column most self-authored postmortems skip. What did the terminal output say? Did the timeout message distinguish between “the command did not run” and “the command ran but did not return”? Was the runbook step ambiguous about whether a retry was safe? Did the deploy tool have a lock, a confirmation prompt, or an idempotency key? If the answer is “the system showed nothing that would have stopped me,” write that down. It is a finding.

Column three: what change would have made the correct action the default. This is where action items come from. The workbook states the principle directly: “In general, trying to change human behavior is less reliable than changing automated systems and processes.” An action item that says “be more careful with migrations” is not an action item. An action item that says “add a preflight check to the migration script that refuses to run if the target table already has a lock row from a prior attempt” is one.

Only column three produces work that survives a personnel change. Columns one and two are the record; column three is the fix.

The template sections, in order

The SRE book’s definition gives the required sections: incident record, impact, actions taken to mitigate or resolve, root cause(s), and follow-up actions to prevent recurrence. The workbook’s case study adds specific failure modes to avoid in each.

Incident record and context

Include a background or glossary section if your service uses internal names that a new hire would not know. The workbook warns: “If you don’t properly contextualize content when writing a postmortem, the document might be misunderstood or even ignored. It’s important to remember that your audience extends beyond the immediate team.” On a small team, the audience beyond the immediate team is the person you hire next quarter.

Impact

Put numbers in it. The workbook is direct: “For outages affecting multiple services, you should present numbers to give a consistent representation of impact… Even if there is no concrete data, a well-informed estimate is better than no data at all.” For a lean team, that might be “approximately 40 minutes of elevated 5xx rate on the checkout endpoint, affecting an estimated 3% of requests during the window.” The estimate is better than the absence.

Actions taken to mitigate or resolve

Do not leave this empty. The workbook’s bad-postmortem example left the recovery efforts section blank and flagged it as a key omission. Write what you actually did: rolled back the deploy, restarted the worker pool, failed over the read replica. This section is also where you record what worked. The workbook notes that a good postmortem lets readers know “what happened, how the issue was mitigated, and how users were impacted.”

Root causes and trigger

This is where the three-column method does its work. The trigger is the event that started the incident. The contributing causes are the conditions that let the trigger produce the impact. The workbook’s case study criticizes a postmortem that “contains a small paragraph that describes the root causes and trigger, but it doesn’t explore the lower-level details of the issue.” Go lower. If the root cause is “the migration script was not idempotent,” the contributing causes include why the script was written that way, why no test caught it, and why the runbook did not warn against a retry.

Action items

The workbook lists the failure modes precisely. Action items should include preventative items, not only mitigative ones. They should use concrete verbs instead of “improve” or “make better,” because “these terms are vague and open to interpretation” and “make it difficult to measure and understand success criteria.” They should have differentiated priorities, not all P2. They should have a tracking bug, because “without a formal tracking process, action items from postmortems are often forgotten, resulting in outages.” And they should have a single owner: “Ideally, an owner is a single point of contact who is responsible for the postmortem, follow-up, and completion.”

For a team of two to fifteen, the single owner is usually the person who wrote the postmortem. That is fine. The point is that the name is on the item, not that the name is on the incident.

Review when you have no review committee

The SRE book states: “An unreviewed postmortem might as well never have existed.” It describes a review process where teams “share the first postmortem draft internally and solicit a group of senior engineers to assess the draft for completeness,” with criteria including whether key incident data was collected, whether impact assessments are complete, whether the root cause is sufficiently deep, whether the action plan is appropriate, and whether the outcome was shared with relevant stakeholders.

A team of five does not have a group of senior engineers to convene. The lean adaptation, which is a proposal rather than a source finding, is a single named reviewer who was not involved in the incident, plus a short written checklist drawn from those criteria. The reviewer’s job is not to judge the author. It is to ask whether the root cause section goes deep enough and whether the action items have owners and tracking bugs. One reviewer, one pass, one checklist.

Publish promptly. The workbook’s bad-postmortem example was published four months after the incident, and “in the interim, had the incident recurred (which in reality, did happen), team members likely would have forgotten key details that a timely postmortem would have captured.” On a small team, the details you lose in four months are the ones you need most: the exact command, the exact error message, the exact state of the system.

Share as widely as is useful. The workbook states: “The value of a postmortem is proportional to the learning it creates. The more people that can learn from past incidents, the less likely they are to be repeated.” For a lean team, that might mean the whole engineering group, or it might mean the two people who will be on call next month. The point is that the document is not private to the author.

What this does not fix

A blameless postmortem does not make the incident less annoying. It does not change the fact that you were the one who ran the command. It does not, on its own, prevent the next incident. The SRE book describes blameless postmortems as “a tenet of SRE culture” and the workbook describes them as important to “creating and maintaining a successful SRE organization,” but neither source claims that writing one is sufficient. The PagerDuty postmortem guide frames the practice as one that “allow[s] your teams to iteratively improve your infrastructure and incident response process” — iteratively, not instantly.

What it does is keep the record honest. The alternative, on a small team, is a postmortem that either blames the author or omits the author’s actions. The first teaches the team that incidents are personal. The second teaches them nothing. The SRE book’s warning is the one to keep in view: “An atmosphere of blame risks creating a culture in which incidents and issues are swept under the rug, leading to greater risk for the organization.”

Write the actions down. Put the system in the causal chain. Give the action items owners and tracking bugs. Then publish it before you forget the error message.

FAQ

Does a blameless postmortem mean I should not name what I did?

No. The SRE book states that the postmortem format “clearly identifies the actions that led to the incident.” The actions stay. What changes is that the postmortem focuses on contributing causes without indicting the individual, and assumes the person acted with good intentions and the information they had.

What if the only honest root cause is that I ran the command again?

Write that sentence, then write the next one: why was running it again possible, and what would have stopped it? The workbook’s principle is that changing automated systems and processes is more reliable than trying to change human behavior. The second sentence is where the action item lives.

How many action items should a postmortem have?

The sources do not prescribe a number. The workbook’s criteria are about quality: at least one preventative item, concrete verbs, differentiated priorities, a tracking bug, and a single owner per item. A postmortem with three well-owned items is more useful than one with ten vague ones.

Who should review a postmortem on a team with no SRE department?

The sources describe review by senior engineers against a set of completeness criteria. A lean-team substitute, which is a proposal rather than a source finding, is one named reviewer who was not involved in the incident, working from a short checklist: was the data collected, is the impact assessed, is the root cause deep enough, are the action items owned and tracked, and were the relevant people told.

Should I write a postmortem for a near miss?

The SRE book lists common postmortem triggers including user-visible downtime, data loss, on-call intervention, resolution time above a threshold, and monitoring failure. It also notes that “any stakeholder may request a postmortem for an event.” A near miss that would have met one of those triggers if a single condition had differed is a reasonable candidate, but the sources do not prescribe a rule for near misses specifically.