Most teams can state their backup schedule. Fewer can state their recovery point objective as a measured interval. The schedule is a claim; the RPO is the number you get when you actually restore and count how many minutes of committed work disappeared. For a lean team running its own on-call rotation, the gap between those two numbers is where incidents get worse than they needed to be.
This is a procedure for closing that gap. It uses PostgreSQL continuous archiving as the worked example because the mechanics are documented precisely, but the method applies to any system where you can name the last recoverable transaction and compare it to the moment of failure.
RPO is an interval, not a schedule
A nightly base backup at 02:00 tells you when the backup started. It does not tell you how much committed data you lose if the primary dies at 14:37. That number depends on two things: how far the archived write-ahead log (WAL) chain reaches, and how long the restore itself takes. The PostgreSQL documentation is direct about the first: point-in-time recovery lets you restore the database “to its state at any time since your base backup was taken.” The second is wall-clock time you have to measure, not estimate.
So the real RPO is the larger of two intervals:
- Archive lag — the gap between the last committed transaction and the last WAL segment safely stored in archive storage.
- Restore time — the wall-clock duration from “we decide to restore” to “the database is serving reads and writes again.”
If archive lag is 40 minutes and restore takes 25 minutes, your real RPO is not 40 minutes. It is 65 minutes, because the data you can recover is already 40 minutes stale before you start, and you spend another 25 minutes getting it back online.
Why the archive lag is usually worse than you think
PostgreSQL invokes the archive command only on completed WAL segments. A segment is normally 16MB. On a busy system, segments fill quickly and archive lag stays small. On a quiet system — nights, weekends, low-traffic services — a segment can sit partially full for a long time. The documentation states the consequence plainly: “if your server generates only little WAL traffic (or has slack periods where it does so), there could be a long delay between the completion of a transaction and its safe recording in archive storage.”
That delay is your archive lag. It is invisible in normal operation because nothing is failing. It becomes visible the moment you need to restore.
The documented lever is archive_timeout. Setting it forces the server to switch to a new WAL segment at least that often, which bounds how old unarchived data can be. The docs note that a setting of “a minute or so” is usually reasonable, and warn that very short values bloat archive storage because forced-switch segments are still full-length files. You can also force a switch manually with pg_switch_wal when you want a just-finished transaction archived immediately.
There is a second constraint that matters for the lottery: to recover using continuous archiving, you need a continuous sequence of archived WAL files extending back at least as far as the start time of your base backup. If the chain is broken — a missing segment, a failed archive command that was never retried, a retention policy that deleted a segment you still needed — the restore stops at the break. The lottery will find this.
The restore lottery: pick a minute, restore to it, measure the gap
The drill is simple to describe and deliberately inconvenient to run. That is the point.
- Pick a random past minute. Not a convenient one. Not the end of a quiet window. Use a timestamp from the middle of a normal working period, ideally one you did not choose in advance. Write it down before you start.
- Restore the base backup to a scratch environment. This is the file-system-level backup plus the WAL chain. pgBackRest’s user guide defines a restore as requiring “the backup files and one or more WAL segments in order to work correctly,” which is the same shape whether you use pgBackRest,
pg_basebackupplus manual WAL shipping, or another tool. - Replay WAL up to your chosen timestamp. In PostgreSQL this is the
recovery_target_timesetting in the recovery configuration. The server replays archived segments until it reaches the target, then stops. - Record the timestamp of the last recoverable transaction. This is the number that matters. Compare it to your chosen target. The difference is your measured archive lag for that moment.
- Record the wall-clock restore time. From “start restore” to “database accepts connections.” This is the second half of your real RPO.
- Note which segment boundary caused the gap. If the last recoverable transaction is 38 minutes before your target, and the next archived segment starts 38 minutes later, you have found the boundary. That is a configuration fact, not a mystery.
Do this on a schedule — monthly is a reasonable starting cadence for a small team — and vary the target time. A drill that always restores to the same convenient minute proves the backup works. It does not measure the RPO.
A worked example
Suppose a team targets a 5-minute RPO. They run a lottery against a quiet Sunday window and pick 03:17 as the target. The restore completes. The last recoverable transaction is at 02:39. The gap is 38 minutes.
Nothing failed. The backup restored correctly. The WAL chain was intact. But the last archived segment before 03:17 was completed at 02:39, and no further segment filled until 03:22 because the system was idle. The archive command had nothing to archive. The measured RPO for that window is 38 minutes, not 5.
The fix is not a new backup tool. It is archive_timeout = 60s, which forces a segment switch every minute and bounds the lag. The team re-runs the lottery the following week against a different quiet window and measures again. This time the gap is under two minutes. The restore time is 22 minutes. The real RPO is now roughly 24 minutes — still worse than the 5-minute target, but now the team knows why and can decide whether to invest in faster restore or accept the number.
That example is illustrative, not a reported incident. The point is the shape of the finding: the drill, not the backup tool, revealed the gap.
What the lottery measures that a backup test does not
A backup test answers “can I restore?” A restore lottery answers “how much do I lose, and how long does it take?” Those are different questions with different failure modes.
The lottery will surface:
- Archive lag during quiet periods. The completed-segment constraint means low-traffic windows accumulate lag.
archive_timeoutis the documented bound. - Broken WAL chains. If a segment is missing, the restore stops early. The lottery finds the break because it tries to reach a specific timestamp.
- Retention that deletes too aggressively. If your retention policy removes WAL segments or base backups before the next one is verified, the recoverable window shrinks. pgBackRest’s documentation notes that incremental backups depend on prior backups being valid to restore, so a retention policy that removes a full backup can invalidate a chain of incrementals.
- Restore time that nobody measured. Teams often know the backup size but not the restore duration. The lottery produces a wall-clock number.
- Object storage versioning gaps. If backup artifacts live in object storage, versioning behavior affects what you can recover. Google Cloud Storage’s Object Versioning documentation explains that a noncurrent version is retained each time a live object is replaced or deleted, and recommends soft delete over Object Versioning for protection against permanent data loss from accidental or malicious deletion. The lottery should confirm that the version you expect to restore is actually the version you get.
Turning the number into a decision
The measured RPO is only useful if it changes something. Three outcomes are common:
The measured RPO is worse than the target, and the gap is archive lag. Shorten archive_timeout. The PostgreSQL docs note that a minute or so is usually reasonable and that very short values bloat archive storage. Test the tradeoff. Re-run the lottery.
The measured RPO is worse than the target, and the gap is restore time. The fix is usually operational: faster storage, a pre-staged base backup, a documented restore procedure that the on-call engineer can follow without re-deriving it. This is where a written recovery checklist earns its keep — see Write the Recovery Checklist Before You Need It for the shape of that document.
The measured RPO is worse than the target, and the team decides to accept it. That is a legitimate outcome, but it should be written down. “We accept up to 40 minutes of data loss during quiet windows because the cost of eliminating it exceeds the value” is a decision. Leaving it unstated is not.
In all three cases, the measured number and the drill procedure go into the runbook. The next on-call engineer should be able to reproduce the measurement without re-deriving the method. That is how the number stays honest as the system changes.
What the lottery does not tell you
The lottery measures recoverability of the database. It does not measure:
- Whether your application can tolerate the restored state. A database restored to 03:17 may be consistent, but the application may have written external state — queue messages, object storage keys, third-party API calls — that the database does not know about.
- Whether your encryption keys are available. If backup artifacts are encrypted, the restore depends on key access. The lottery should include a key-retrieval step, or it is measuring a restore that cannot happen.
- Whether your monitoring would have told you the archive was lagging. The lottery is a drill; monitoring is the production control. If archive lag can grow to 38 minutes without an alert, the lottery will keep finding the same gap.
These are separate measurements. The lottery is the one that produces the RPO number.
Frequently asked questions
How often should we run a restore lottery?
Monthly is a reasonable starting cadence for a team of 2–15 engineers. The goal is to catch drift: configuration changes, retention policy changes, schema growth that slows restore, and quiet periods that expose archive lag. If your system changes frequently, run it more often. If it is stable, quarterly may be enough — but the first few runs should be close together so you can see whether the number is stable.
Can we run the lottery against production?
No. Restore to a scratch environment. The lottery is a measurement, not a failover. Running it against production converts a drill into an incident.
What if we use restic instead of pgBackRest?
The lottery method applies to any backup system where you can name the last recoverable point. For restic, the snapshot timestamp is the starting point. The restic documentation covers snapshot listing, restore, and repository integrity checking. The measurement is the same: pick a target time, restore, record the timestamp of the last recoverable data, and compare. The difference is that restic snapshots are point-in-time by nature, so the “archive lag” is the interval between snapshots rather than the interval between WAL segments.
What if the restore fails?
That is a successful lottery. A failed restore in a drill is a failed restore you did not have during an incident. Record what failed, fix it, and re-run. The PostgreSQL documentation notes that archive commands should refuse to overwrite pre-existing archive files, and that a nonzero exit status tells PostgreSQL the file was not archived and it will retry periodically. If your archive command has been silently failing, the lottery will find the gap in the WAL chain.
Do we need a dedicated SRE to run this?
No. The procedure is a sequence of documented commands and a stopwatch. The value is in running it on a schedule and writing down the number. A team that already runs its own on-call rotation can add this to the rotation as a monthly task.
The number you can defend
A backup schedule is a hope. A measured RPO is a number you can put in a runbook, compare to a target, and defend in a postmortem. The restore lottery is how you get it. Pick a minute, restore to it, and count what you lost. Then decide whether to change the configuration, change the procedure, or accept the number. All three are better than not knowing.