Consistent Database Backups With Btrfs Snapshots + Backrest Hooks
When you back up a directory with restic or Backrest, and that directory contains live databases like PostgreSQL or SQLite, copying the data files directly can leave you with a backup that won’t start — or worse, one that starts silently with corrupt data. This article shows how to freeze the whole data directory into a consistent state with a btrfs subvolume snapshot, and have Backrest create the snapshot automatically via a hook before every backup. The result is in-place backup and in-place restore.
Why file-level backups are unreliable
PostgreSQL: you get a “half crash scene”
The PostgreSQL docs (25.2 File System Level Backup) are explicit: when you copy a data directory while the server is running, recovery relies on replaying WAL (write-ahead log), which requires the WAL to be included in the backup. If your backup config’s excludes contain **/pg_wal/*, the snapshot is missing WAL segments after the last checkpoint, and PG will PANIC on restore (could not locate a valid checkpoint record) rather than “maybe lose some data”.
Even without the WAL exclusion, restic’s online copy has a fundamental flaw: files are read one by one, each at a different point in time. The data files are read at T1, the WAL at T2 (minutes may pass in between while PG keeps writing and checkpoints keep advancing), so the snapshot contains “T1 data files + T2 WAL” — a state that never existed on disk. PG cannot recover from that.
Page tearing and silent corruption
When restic reads an 8KB page while PG is writing it, it may capture a “half old, half new” page:
- With checksums on (
data_checksums=on): detected at startup, reportsinvalid page in block, fails loudly (just restore an earlier snapshot) - With checksums off (default): starts up silently, but that page’s data is corrupt, with no error at all
SQLite in WAL mode
Copying only the SQLite main database file loses recent writes that haven’t been checkpointed from the WAL; restic interleaves reads of the main file and the WAL at inconsistent points in time. The SQLite docs explicitly say copying a live database with cp is unsafe.
A file-level backup of running databases is “survival by probability”: multiple snapshots cover for each other, but it is not a design guarantee. Consistency must come from the database’s own tools or a filesystem snapshot; restic only handles transfer, dedup, and encryption.
Comparing solutions
| Option | Offsite restore | Consistency | Effort |
|---|---|---|---|
| Raw file-level copy | In-place, but may PANIC / corrupt silently | Probabilistic | 0 |
| pg_dump + systemd timer | pg_restore, not in-place | 100% | script + timer |
| btrfs snapshot (this article) | In-place | 100% (btrfs atomicity) | move data dir to subvolume |
If you require in-place backup, in-place restore — back up the data directory itself, and on restore lay it back in place and start it directly, no dump/import round-trip — pg_dump is not a fit (restore requires pg_restore). The btrfs snapshot approach is the most direct path.
Why btrfs snapshots guarantee consistency
A btrfs snapshot is not a “file-level backup” — it is a CoW (copy-on-write) metadata operation that freezes the entire subvolume in milliseconds. Every file in the snapshot (data files, WAL, pg_control) comes from the same instant T0. This is exactly the “consistent snapshot” model PostgreSQL recognizes:
A snapshot is the disk state of a sudden power loss; WAL is written to disk before data (write-ahead), so the WAL in the snapshot always covers every transaction the data files reflect; on recovery PG replays WAL from the checkpoint — the same path it takes on every boot after a crash, fully mature.
The key difference: restic’s online copy reads files at random instants (T1/T2/T3), while a btrfs snapshot captures everything at T0. PG recovery only needs the one guarantee that “the whole tree comes from a single instant”.
Fixed-name snapshot + rotation (key design)
You don’t need to pile up historical snapshots locally — the restic snapshot chain IS the history (blob dedup stores only deltas). Locally, only one fixed-name snapshot is needed, deleted and recreated before every backup:
Every 3h (backrest SNAPSHOT_START hook):
1. btrfs subvolume delete snapshots/data-latest # delete old (seconds)
2. btrfs subvolume snapshot -r data snapshots/data-latest # freeze current
backrest: back up /path/to/snapshots/data-latest # path never changes
→ restic compares the same path → uploads only blocks changed in the 3h window (incremental)
→ each restic snapshot = a complete consistent state at that time
→ retention 24h + 30d + 12m = 66 consistent versionsHistory is carried entirely by the restic snapshot chain; locally only data-latest remains as the “quiescent source”.
Prerequisite: move the data directory to its own btrfs subvolume
btrfs subvolume snapshot only works on subvolumes. If your data directory (this article uses ~/data as an example) is a plain directory, move it to its own subvolume first. The migration requires stopping all containers, so script it:
# 1. Stop all container services
# 2. Create the new subvolume + fix ownership + reflink copy (CoW, no double space)
sudo btrfs subvolume create /path/to/data.new
sudo chown $USER:$USER /path/to/data.new
cp -a --reflink=always /path/to/data/. /path/to/data.new/
# 3. Consistency check (file count + rsync per-file comparison)
[[ $(find /path/to/data -type f | wc -l) -eq $(find /path/to/data.new -type f | wc -l) ]] || exit 1
rsync -a --dry-run --itemize-changes /path/to/data/ /path/to/data.new/ | grep . && exit 1
# 4. Atomic rename (mv only changes the path entry; subvolume identity/subvolid stays)
mv /path/to/data /path/to/data.old
mv /path/to/data.new /path/to/data
# 5. Restart all containersVerify subvolume identity: sudo btrfs subvolume show /path/to/data showing a Subvolume ID means success. Keep data.old for a few days, then sudo rm -rf /path/to/data.old once confirmed.
Write the snapshot script
~/.local/bin/data-snapshot.sh, fixed-name rotation with a re-entry lock:
#!/usr/bin/env bash
set -euo pipefail
DATA=/path/to/data
SNAP_DIR=/path/to/snapshots
SNAP="$SNAP_DIR/data-latest"
# Re-entry lock
if [[ -f /tmp/data-snapshot.lock ]]; then
kill -0 "$(cat /tmp/data-snapshot.lock)" 2>/dev/null && { echo "already running"; exit 1; }
rm -f /tmp/data-snapshot.lock
fi
echo $$ > /tmp/data-snapshot.lock
trap 'rm -f /tmp/data-snapshot.lock' EXIT
mkdir -p "$SNAP_DIR"
[[ -d "$SNAP" ]] && sudo btrfs subvolume delete "$SNAP"
sudo btrfs subvolume snapshot -r "$POD" "$SNAP"
[[ -d "$SNAP" ]] || exit 1If sudo is not NOPASSWD, allow only these two fixed commands in sudoers.
Configure Backrest: point paths at the snapshot + hook
Edit ~/data/backrest/config/config.json (back it up first):
{
"plans": [
{
"id": "remote",
"repo": "remote",
"paths": ["/path/to/snapshots/data-latest"],
"excludes": [
"*.log",
"**/cache/*",
"**/tmp/*",
"**/pg_xlog/*",
"**/valkey/*",
"**/redis/*",
"**/kuma.db*",
"..."
],
"hooks": [
{
"conditions": ["CONDITION_SNAPSHOT_START"],
"onError": "ON_ERROR_FATAL",
"actionCommand": {
"command": "/path/to/.local/bin/data-snapshot.sh"
}
}
],
"schedule": { "maxFrequencyHours": 3, "clock": "CLOCK_LAST_RUN_TIME" },
"retention": {
"policyTimeBucketed": { "hourly": 24, "daily": 30, "monthly": 12 }
}
}
]
}Trigger a backup to verify the full chain
# Connect RPC endpoint (not the bare /v1/backup, 405 is misleading)
curl -X POST http://127.0.0.1:9898/v1.Backrest/Backup \
-H 'Content-Type: application/json' -d '{"value":"remote"}'Verification points: journalctl shows the hook’s sudo btrfs records → restic process starts → restic snapshots shows a new snapshot (tag plan:remote, --parent correctly reuses the old chain for incremental).
Restore (in-place)
Local restore (disk intact, seconds)
# Copy the snapshot back directly
sudo btrfs subvolume snapshot /path/to/snapshots/data-latest /path/to/data.restored
# or restore from any historical snapshotOffsite restore (disk dead, pull from restic)
# 1. Restore the snapshot directory from restic (pick any consistent version)
restic restore <snapshot-id> --target /path/to/restore/
# 2. Recreate the subvolume and lay it back
sudo btrfs subvolume create /path/to/data
cp -a --reflink=always /path/to/restore/snapshots/data-latest/. /path/to/data/
# 3. Start the containers — PG replays WAL via standard crash recovery, ready to useWhy in-place restore works: the snapshot is a perfect crash scene; PG’s recovery path is identical to a power-loss reboot — no pg_restore, no cluster rebuild.
References
- PostgreSQL docs 25.2 File System Level Backup: copying a running server requires complete WAL, or use a filesystem consistent snapshot
- SQLite Backup API docs: copying a live database with cp is unsafe; use the Online Backup API / VACUUM INTO
- restic docs 040_backup:
--stdin-from-commandis better than a pipe, avoids masking errors - Backrest (WebUI orchestrator for restic) source:
gen/go/v1/config.pb.go(hook structure),internal/env/environment.go(env vars),gen/go/v1/service_grpc.pb.go(RPC endpoints)