Point-in-Time Recovery (PITR)
Point-in-Time Recovery — Restore to Any Second in the Retention Window
UVP
When the dropped table happened at 14:32:07 and the backup ran at 02:00, you don’t want to lose 12 hours of data — you want to recover to 14:32:06. The Full edition’s PITR coordinator combines periodic snapshots with continuous WAL archival so you can restore to any LSN, any timestamp, any transaction ID, or any named recovery point inside the retention window. Configurable RPO down to 1 minute, RTO down to 5 minutes, parallel recovery workers, checksum verification, optional compression. No external backup service. No vendor lock-in. Local, S3, Azure, GCS — pick your archive directory.
Prerequisites
- A running HeliosDB Full instance.
- Disk space for the WAL + snapshot + archive directories.
- Optional: object-storage destination if you don’t trust local disk.
- About 20 minutes.
1. The Configuration
RPO (max acceptable data loss) can be one minute, five minutes, fifteen minutes or one hour; RTO (max acceptable recovery time) can be five, fifteen or thirty minutes, or one hour.
The defaults give you RPO=1min, RTO=30min, hourly snapshots, 16MB WAL segments, gzip-6 archives. Fine for most production. Tighten as needed; nothing else changes.
2. Continuous Archiving
Once PITR is initialized, the system is archiving WAL continuously (subject to wal_segment_size_mb rollover) and creating snapshots every snapshot_interval_secs.
3. The Four Recovery Targets
- LSN — recover to a specific WAL log sequence number
- Timestamp — recover to a specific wall-clock time
- Transaction — recover to just before / just after a txid
- Latest — most recent recoverable point
- Recovery point — a previously named point
Timestamp is the one you’ll use most. The others are forensic.
4. Recover to a Timestamp
The recovery engine:
- Finds the most recent snapshot at or before the target time.
- Restores the snapshot to
target_directory. - Replays WAL records from the snapshot LSN up to the target timestamp.
- Verifies checksums on every WAL record (when checksum verification is enabled).
- Returns once the target has been reached — earlier than the snapshot is unreachable, later requires fresh WAL.
It does not overwrite your live database. Recovery goes to a fresh target_directory. You promote it manually when you’re ready.
5. Recover Just One Table
Partial recovery mode tells the engine to skip WAL records that don’t touch the listed tables. Useful when only one table got nuked and you don’t want to replay everyone else’s last 12 hours of writes.
6. Validate Without Recovering
Validation-only mode walks the WAL chain, verifies checksums, and confirms the target is reachable — without writing anything. Run this in your DR drill cron job to make sure your archives are actually usable. If checksums fail, you find out before the disaster.
7. Storage Backends
PITR supports S3, Azure, GCS, and local archive destinations.
Set archive_directory to a path that your storage layer maps to your bucket:
| Destination | archive_directory example |
|---|---|
| Local disk | /var/lib/heliosdb/archive |
| S3 | s3://my-bucket/heliosdb/pitr |
| Azure Blob | az://account/container/heliosdb |
| GCS | gs://my-bucket/heliosdb/pitr |
Compression (enable_compression: true) and archive compression level (0–9) are independent of the destination.
8. RPO/RTO Tuning
| Goal | RPO | RTO | snapshot_interval_secs | recovery_workers |
|---|---|---|---|---|
| Compliance baseline | OneHour | OneHour | 3600 | 4 |
| Standard prod | OneMinute | ThirtyMinutes | 3600 | 4 |
| Tight SLA | OneMinute | FiveMinutes | 600 | 8 |
Tightening RPO costs WAL archive bandwidth. Tightening RTO costs more snapshots (so there’s less WAL to replay) and more recovery workers.
The PITR engine raises an RTO-exceeded error (with the expected and actual durations) if a recovery overruns its target — useful for SLA monitoring.
9. Named Recovery Points
Before a risky deploy, mark a known-good named recovery point (e.g. pre-v8.0.3-release); if the deploy goes wrong, recover to that point.
Named points survive WAL truncation as long as their underlying LSN is still inside max_wal_segments.
10. Plug into the Coordinator from SQL
Recovery is a binary-level operation; the SQL surface is for inspection:
-- See available recovery pointsSELECT id, timestamp, wal_lsn, snapshot_id, size_bytesFROM pg_recovery_pointsORDER BY timestamp DESC LIMIT 20;
-- See current WAL positionSELECT pg_current_wal_lsn(), pg_last_wal_replay_lsn();Recovery itself is run from the coordinator binary — never from the live SQL session you’re trying to roll back.
How to Configure
Scheduled backups, WAL archiving, point-in-time restore, the RPO/RTO targets, the snapshot interval, recovery workers and named recovery points are not set through a public configuration file or command-line interface. Configuration for this feature is provided during onboarding — contact support@heliosdb.com (or sales@heliosdb.com if you are not yet a customer).
The heliosdb.toml file of the HeliosDB Full server accepts only the documented storage keys (storage.data_dir, storage.memtable_size_mb, storage.compaction_strategy, storage.read_cache_mb and the storage.prefetch_* settings). The server refuses to start if the file contains any other key.
To check backup state from SQL:
SHOW BACKUP STATUS;Where Next
- raft-setup.md — PITR rides on top of the per-node WAL.