Rebuilding an Incident Timeline from Slack, Logs, and PagerDuty When Nobody Took Notes

Nobody scribed. That is the normal case, not the failure case. The incident channel has 400 messages, the alert history has a dozen state transitions, and the application logs have a few thousand lines. The postmortem is due in three days. What you have is not a memory problem; it is a reconstruction problem, and reconstruction is a procedure.

The useful framing is that you are reconciling three independent clocks. PagerDuty records when alerts fired, were acknowledged, and resolved. Your log platform records when the system did things. Slack records when people said things. Only two of those three are reliable for ordering events, and the third is the only source for intent. Treating them as interchangeable is what produces timelines that look precise and are wrong.

Name an owner before the channel goes quiet

The single highest-leverage move for a 2–15 engineer team is not a new tool. It is designating a postmortem owner at incident close, while the responders are still in the channel. PagerDuty’s public postmortem process describes this role explicitly: the owner is responsible for populating the postmortem, looking up logs, managing the follow-up investigation, and keeping interested parties in the loop (PagerDuty, Postmortem Process).

That sentence converts “nobody took notes” from a memory failure into a scheduled task. The incident commander can name the owner in the final minutes of the call. The owner does not need to have been the deepest responder; they need to be the person who will spend the next afternoon in query consoles and Slack search.

Google’s SRE book frames the same artifact more broadly: a postmortem is a written record of an incident, its impact, the actions taken to mitigate or resolve it, the root cause(s), and the follow-up actions to prevent recurrence (Google SRE Book, Chapter 15). The timeline is the spine that holds those sections together. Without it, the analysis floats free of the events.

Set the reconstruction window and pick the authoritative clock

Before opening any tool, write down the window. Start it earlier than the first alert — the causal chain usually begins before the first page — and end it after the last state transition, not after the last Slack message. A window that starts at the first alert and ends at the last chat message will systematically miss the change that caused the incident and the cleanup that followed it.

Then decide which clock governs which class of event:

  • Alert and escalation transitions: the monitoring platform’s timestamps. These are machine-generated and monotonic within the platform.
  • System behavior: log timestamps, with the caveat that ingestion delay and clock skew are real. Note the offset between the log platform’s receive time and the event’s own timestamp field where both exist.
  • Human decisions: Slack message times, treated as approximate. They tell you ordering within a conversation, not absolute time.
  • Database recovery events: PostgreSQL Log Sequence Numbers, which increase monotonically with each WAL record and can be compared to measure the volume of WAL data between two points (PostgreSQL Documentation, WAL Internals). LSNs order database events precisely. They say nothing about human decisions, and you should not ask them to.

Writing this mapping down once, in the incident channel or the postmortem draft, prevents the most common reconstruction error: quoting a Slack timestamp as if it were an event time.

Pull the machine record first

Query the log platform before you read a single chat message. The machine record gives you a monotonic spine, and reading Slack first anchors your reconstruction on the least reliable clock.

If you are on CloudWatch Logs Insights, the query language supports the operations you need for this pass. The filter command returns only log events matching one or more conditions. The sort command orders results ascending or descending. The limit command caps the number of events returned, which is useful with sort to get the most recent or top results rather than the full window (AWS, CloudWatch Logs Insights query syntax).

Two commands are worth knowing for reconstruction specifically. join combines log events from a source log group with events from another log group or query result based on a matching field, which is how you correlate related events across sources using a request identifier or transaction ID. sessionize groups events into sessions by identity fields and an inactivity gap, which is how you turn a stream of retries into a small number of logical attempts. Both are documented in the same reference.

If the incident involved a database recovery, pull the backup tool’s logs next. pgBackRest’s user guide notes that a restore requires the backup files and one or more WAL segments to work correctly, and that WAL segments follow a naming convention where the first eight hexadecimal digits represent the timeline and the next sixteen are the logical sequence number (pgBackRest User Guide). Those names are themselves a timeline. If you restored from a restic repository, the troubleshooting section of the restic documentation covers finding damaged data, backing up the repository, repairing the index, and checking the repository again (restic Documentation) — useful when the reconstruction reveals that the restore itself had a gap.

One operational note worth carrying into the timeline: pgBackRest requires that local and remote versions match exactly, and a mismatch stops WAL archiving and backups from functioning until versions match. If your incident involved a failed archive, the version mismatch is a candidate cause and belongs in the timeline as an observed fact, not a hypothesis.

Mine Slack second, and read it narrowly

Slack is the only source for intent. It tells you who decided what, when a hypothesis was abandoned, and which workaround was tried before the one that worked. It is a poor source for what happened when.

PagerDuty’s process is direct about this use: go through the history in Slack to identify the responders and add them to the page. Identify the incident commander and scribe in that list. The same process treats the timeline as the main focus when beginning to populate the postmortem, and specifies that it should include important changes in status or impact and key actions taken by responders.

Read Slack with a purpose. You are looking for four things:

  1. Decision points. “Rolling back the deploy” is a timeline entry. The discussion that preceded it is context.
  2. Abandoned hypotheses. These are the most valuable and most commonly lost entries. A hypothesis that was tested and rejected is a fact about the incident, and it prevents the next responder from re-testing it.
  3. Presence. Who joined the channel, when, and from which team. This is the map of who can corroborate which rows.
  4. Workarounds. What was tried, in what order, and what the observed effect was.

Do not read Slack for timestamps. Read it for content, then attach the nearest machine timestamp as the row’s time and mark the row as approximate.

Reconcile into one table with a source column

The output of the reconstruction is a table, not prose. Each row has a time, an event, a source, and a flag for whether the row is observed or inferred.

A minimal schema:

Time Event Source Observed / Inferred
T+00:00 First alert fires PagerDuty alert history Observed
T+00:14 Escalation to secondary PagerDuty escalation history Observed
~T+00:22 Failover initiated Slack #inc-1234, time approximate; corroborated by PagerDuty escalation at T+00:14 Inferred
T+00:31 Error rate returns to baseline CloudWatch Logs Insights query, permalink in appendix Observed

The third row is the important one. A hypothetical example, clearly labeled as such: a team discovers at postmortem time that the only record of a failover decision is a one-line Slack message with no timestamp context. Under this approach the row reads as above — approximate time, named source, corroborating machine event — rather than a confident but unsourced time. A visible gap is more useful to a reviewer than a plausible sentence.

PagerDuty’s process asks for exactly this discipline: for each item in the timeline, identify a metric or some third-party page where the data came from — a graph link, a search, a post — anything that shows the data point you are trying to illustrate. It also asks that any commands or queries used to look up data be posted on the page so others can see how the data was gathered. That second requirement is what makes the timeline reviewable rather than merely readable.

Write each row so a reviewer can re-derive it

A timeline entry that cannot be re-derived is an assertion. The test is simple: can a reviewer who was not on the call take the row’s source, run the query or open the permalink, and see the same thing?

For log-derived rows, paste the query. For alert-derived rows, link the incident in the monitoring platform. For Slack-derived rows, link the message permalink. For database recovery rows, cite the LSN range or the WAL segment names. The cost of this discipline is a few minutes per row. The benefit is that the postmortem meeting does not spend its first fifteen minutes arguing about whether an event happened at T+22 or T+31.

Google’s SRE book is blunt about the alternative: an unreviewed postmortem might as well never have existed. The review is what converts a document into a shared understanding, and a timeline without sources cannot be reviewed — only believed or disbelieved.

Circulate the draft before the meeting

PagerDuty’s process recommends posting a link to the postmortem into Slack for internal review of style and content roughly 24 hours before the meeting, so that experienced readers can flag missing detail while there is still time to fix it. For a small team, this is the difference between a meeting that resolves disagreements and a meeting that discovers them.

The review pass has a specific job for the timeline: check that every row’s source actually supports the row. This is where inferred rows get promoted to observed, or get marked as gaps. It is also where a reviewer who was on the call can supply a missing corroborating source — a graph they had open, a query they ran — that the owner could not have known about.

The meeting itself, per the same process, opens by recapping the timeline to make sure everyone agrees and is on the same page. That recap is short when the timeline is sourced and long when it is not.

Keep the recipe as a runbook

The reconstruction procedure — the window definition, the clock mapping, the query patterns, the table schema, the review window — is itself a runbook artifact. Written down once, the next incident’s timeline costs an afternoon instead of a week of intermittent effort. It can also be rehearsed: a pre-mortem that walks through the reconstruction steps on a past incident surfaces the gaps in your logging and your chat conventions before you need them under pressure.

This connects to a broader practice worth adopting: writing the recovery checklist before you need it. The same logic applies to postmortem reconstruction. The checklist you write during a calm week is the one you can follow during a bad one.

Two conventions make the recipe durable. First, keep the timeline table in the postmortem template, with the source and observed/inferred columns already present, so the owner does not have to invent the schema under time pressure. Second, keep a short appendix of the queries used, with the incident window parameterized, so the next owner can adapt rather than rediscover.

What this does not fix

Reconstruction is not a substitute for contemporaneous notes. A scribe who captures decisions in real time produces a better timeline than any amount of after-the-fact querying. But a small team without a dedicated SRE function will not always have a scribe, and the reconstruction procedure is what keeps that from becoming a postmortem that never gets written.

It also does not fix missing data. If the log retention window closed before the postmortem owner started, the row is a gap. Label it as a gap and move on. The gap itself is a finding — it tells you something about retention policy that is worth a follow-up ticket.

Finally, the procedure does not make the timeline blameless by itself. Blamelessness is a property of how the rows are written and how the meeting is run. Google’s SRE book defines a blameless postmortem as one that focuses on identifying the contributing causes of the incident without indicting any individual or team for bad or inappropriate behavior. A timeline that records “engineer X ran the wrong command” is not blameless; one that records “the runbook’s rollback step assumed a different deploy mechanism than the one in use” is.

FAQ

What if the incident channel was archived or the messages are gone?

Then Slack is not a source for that incident, and the timeline is built from the machine record alone. Mark the human-decision rows as gaps. The reconstruction procedure still works; it just produces a thinner timeline. The follow-up action is a retention or export decision, not a reconstruction technique.

How do I handle timezone differences between responders?

Normalize every timestamp to UTC in the table, and note the original timezone in the source column if it matters. The machine sources — alert history, log timestamps — are typically already UTC or carry an explicit offset. Slack displays in the viewer’s local time, which is one more reason to treat its timestamps as approximate.

Should the timeline include events that turned out to be irrelevant?

Yes, if they consumed responder attention. An abandoned hypothesis is a fact about the incident. The timeline’s job is to record what happened, including the dead ends, so the next responder does not re-walk them.

How detailed should each row be?

Detailed enough that a reviewer can re-derive it from the cited source. If the source is a query, the row needs the query. If the source is a graph, the row needs the link and the time range. If the source is a Slack message, the row needs the permalink and a note that the time is approximate.

What if the postmortem owner was also a responder?

That is common on small teams and it is workable, with one caveat: the owner’s own actions are the hardest rows to source, because the owner remembers them and may not have written them down. Flag those rows explicitly and ask a second responder to corroborate them during the review pass.

Does this work for a database recovery incident specifically?

Yes, and the database gives you better anchors than most incidents. PostgreSQL LSNs are monotonic and comparable, so you can order recovery events precisely. pgBackRest WAL segment names encode the timeline and logical sequence number, so the archive history is itself a timeline. The human decisions around the recovery — when to fail over, when to accept data loss — still come from Slack and still need the approximate-time treatment.

The short version

Name a postmortem owner at incident close. Define the window and the authoritative clock for each event class. Pull the machine record first, Slack second. Reconcile into a table with a source column and an observed/inferred flag. Cite a query, graph, or permalink for every row. Circulate the draft a day before the meeting. Keep the recipe as a runbook. The timeline that results is not a reconstruction of memory; it is a reconstruction of evidence, and it is reviewable by someone who was not there.