Backups and Restore
FreeSDN ships backup mechanisms with different scopes. Using the wrong one during an incident is costly - understand each before you need them.
Backup mechanisms
Section titled “Backup mechanisms”| Mechanism | What it captures | Use case |
|---|---|---|
| pg-backup (DB dumps) | Application and LogDB database contents, including stored encrypted credentials; not application file volumes, broker state, arbitrary roles/ACLs or external device data | Database recovery as one part of a coordinated full-instance recovery |
Configuration Backup → .fsdn |
Config snapshot: Sites, Controllers, Devices, users, automation. No secrets; tied to the source instance key | Config portability, version history, dev-to-prod config copy |
Configuration Backup → .fsdnvault |
Supported configuration contributors and their secrets, sealed under an operator passphrase and re-keyed on restore; not every database row, volume, queue or external VM/disk | Guided configuration and credential recovery into a compatible empty target |
pg-backup: daily GPG-encrypted DB dumps
Section titled “pg-backup: daily GPG-encrypted DB dumps”The freesdn-pg-backup container is part of the core stack - no profile required. It dumps both databases daily:
freesdn- the primary PostgreSQL DB (19 schemas: access, agents, ai, analytics, audit, backup, cameras, collector, core, devices, enterprise, events, fabric, firewall, gateway, hypervisor, network, voip, vpn)freesdn_logs- the TimescaleDB observability DB
Dumps land in the pg_backups volume with 7-day local retention.
GPG encryption (required in production)
Section titled “GPG encryption (required in production)”By default, pg-backup fail-closes without GPG configured: it refuses to write unencrypted dumps. This applies to every deployment tier - the guard checks only whether GPG is active, not ENVIRONMENT. Any stack that lacks a valid GPG recipient must set BACKUP_ALLOW_PLAINTEXT=1 explicitly (the homelab .env.lite template includes a commented-out opt-in; uncomment it only if you have no GPG key and understand the risk - pro and max tier templates include GPG placeholder values that you must fill in before running - generate a dedicated keypair, set BACKUP_GPG_RECIPIENT to the real recipient email, and place the exported public key at ./secrets/backup-public.asc - and must not set this flag). Configure a GPG recipient before running in production.
Setup:
-
On a secure restore host (not the production server), generate a dedicated backup key pair:
Terminal window gpg --full-generate-key# Choose RSA 4096, no expirygpg --export --armor backup@example.com > backup-public.asc -
Copy
backup-public.ascto./secrets/backup-public.ascin the repo. -
Set in your env file:
Terminal window BACKUP_GPG_RECIPIENT=backup@example.comBACKUP_GPG_PUBLIC_KEY_PATH=./secrets/backup-public.asc
The container holds only the public key and cannot decrypt its own dumps. Store the private key and its passphrase in a secrets manager completely separate from the production server.
Completion health and monitoring
Section titled “Completion health and monitoring”The pg-backup healthcheck requires an enabled backup service, no recorded
failure on its latest attempt, and a completed backup pair whose timestamp is
within BACKUP_INTERVAL_SECONDS + BACKUP_FRESHNESS_GRACE_SECONDS. The defaults
are 24 hours plus 2 hours. Both the application and LogDB final files must exist
and be nonempty. A future, malformed or mismatched completion timestamp fails
health. Old files or a recently created backup directory do not establish health.
The service atomically replaces its local status file at
/backups/.metrics/pg_backup.prom. A failed dump retains the previous success
time but sets freesdn_pg_backup_last_attempt_failed to 1. Disabled encryption
configuration sets freesdn_pg_backup_enabled to 0 and is unhealthy; the service
still idles without writing plaintext. The completion pointer at
/backups/.metrics/pg_backup.last_success identifies the pair covered by that
success time. A completed dump is not evidence that a restore will succeed.
Enable the metrics profile to run the internal backup-metrics exporter and
Prometheus scrape job freesdn-backups. The exporter reads both volumes without
write permission, runs without root or Linux capabilities, and publishes no host
port. Only /metrics and /healthz are served; archive contents, paths and
credentials are not exposed. /healthz means the exporter is reachable, not that
backups are healthy.
The shipped alerts cover missing exporter observations, malformed status, missing archive pairs, disabled backups, failed attempts, stale completion, explicitly unencrypted backups, and required off-site status. Freshness thresholds follow the configured intervals. Missing status produces explicit zero-valued validity metrics rather than disappearing time series. Alerts generally require 5 minutes of continuous failure; stale/retention alerts use 15 minutes and the plaintext warning uses 30 minutes. First runs that have not completed yet are reported as unavailable; large initial dumps can therefore alert before finishing.
Alert delivery is not configured automatically. Prometheus evaluates these rules, but an operator must configure Alertmanager and a notification receiver to receive pages or email. Test that delivery path before relying on it.
The local completion timestamp is not a worker heartbeat. If a worker stops after a successful cycle, age-based alerts fire when the configured limit is exceeded. A completed dump or verified remote copy still does not prove that a restore will succeed.
Manually trigger a backup
Section titled “Manually trigger a backup”The configured service begins a backup cycle when it starts. With a verified GPG
recipient and public key, restart only pg-backup and wait for the new completed
pair before proceeding:
docker compose --env-file .env.pro restart pg-backupdocker compose --env-file .env.pro ps pg-backupdocker compose --env-file .env.pro logs --tail=50 pg-backupUse the selected installation’s actual environment file and project. A zero exit
from restart does not mean the backup finished. Require fresh completion metadata
and nonempty files for both databases, verify encryption and retained keys,
then rehearse restoring the pair. Treat raw logs and file paths as private.
Do not run a separate pg_dump | gzip shell to bypass the backup entrypoint:
that path does not inherit the entrypoint’s encryption checks and can create
plaintext dumps in a production backup volume. A consistent recovery point across
the databases, files and broker also requires stopping/accounting for writers;
two separately completed daily dumps do not establish that global consistency.
List and retrieve backups
Section titled “List and retrieve backups”List files and read the completion pointer privately to select an actual completed pair. Do not infer the database filename from a hardcoded example date. Copy files as binary data, without allocating a TTY:
docker compose --env-file .env.pro exec -T pg-backup ls -lh /backupsdocker compose --env-file .env.pro exec -T pg-backup cat /backups/.metrics/pg_backup.last_successUse docker compose cp with each reviewed exact filename into a protected local
directory, or a non-TTY binary stream with exec -T. Verify the copied bytes against
the source and retain the pair, matching vault configuration and required application
files together. Never publish dumps, secret configuration or private decryption keys
in an issue or release artifact. Test recovery using a separately retained private
key; keeping only the public encryption key is insufficient.
Off-site DR via rclone (dr profile)
Section titled “Off-site DR via rclone (dr profile)”Enable dr and metrics in COMPOSE_PROFILES, and configure a dedicated rclone
remote. See Compose Profiles for setup steps. Set
BACKUP_OFFSITE_REQUIRED=1 to require remote completion evidence even before the
first run, especially if using --profile dr on the command line. The exporter
also infers this requirement from COMPOSE_PROFILES=...,dr or an existing remote
status file. Removing dr after it has been used does not silently clear its
monitoring requirement.
The off-site service copies retained, nonempty database pairs with generated
filenames. It excludes .partial files, orphaned dumps, and manually named files.
It snapshots the local completion pointer and requires both files of that pair.
After copying, it downloads remote contents for comparison and checks that every
selected source file appears in the verified match list. Empty or vanished source
files cannot produce a false completion. Once verified, it writes and reads back
_completed/<batch>.<extension>.txt on the remote, then atomically records status
in the separate pg_backup_offsite_status volume. Archives remain mounted read-only.
Verification reads all selected retained backup data on each cycle. Budget for network bandwidth, provider requests, egress charges and completion time. A failed copy, content check or marker operation preserves the previous success and records a failed latest attempt. Rechecking the same old backup updates the verification time but does not reset its source age. These observations apply to the last verification; remote loss after that check is detected on a later run.
Off-site retention defaults to 30 days, independently of local retention. The
currently acknowledged pair and marker are protected from that sweep, even when
old. Retention errors are reported separately and fail container health without
erasing successful copy evidence. The rclone --immutable option prevents this
client from replacing different existing content; it is not storage-enforced
immutability. Configure object lock/versioning, separate credentials and an
independent retention policy with your storage provider. Do not weaken object-lock
protection merely to silence a deletion warning.
Recovery points depend on both schedules. Daily local dumps can leave about 24 hours of data since the previous snapshot. A separate daily off-site copy loop can add nearly another 24 hours, plus dump/transfer/verification time. Neither interval is a guaranteed RPO or RTO. With defaults, remote completion-age alerts use 24h local + 24h copy + 2h grace. Use shorter intervals or a separately tested continuous backup design when the recovery objective is tighter. Measure it with restores from the actual remote storage and keep the decryption keys off the production server.
The Configuration Backup module
Section titled “The Configuration Backup module”The Configuration Backup module (at /backup in the UI) produces portable .fsdn archives. Use them to:
- Migrate configuration from a staging instance to production
- Snapshot configuration before a major change
- Seed a new Site with an existing Site’s configuration
A .fsdn archive carries no credentials. After restoring a config archive to a new instance, re-enter all module secrets (Controller passwords, NVR credentials, etc.).
To move supported configuration and its secrets, the module produces a .fsdnvault archive, sealed under an operator passphrase and re-keyed using the destination’s encryption configuration on restore. It also drives the setup wizard’s “Restore from a backup instead” branch on a compatible fresh instance. Restored organization/admin configuration must pass login, credential, contributor and integration checks. This archive does not restore external VM disks, media, broker state or every database row and application file. See Configuration Backup → secure full backup for the create/restore flow and credential-handling requirements.
The module supports selective restore (individual sections), strict semver schema gating (a mis-versioned section is rejected, not silently applied), and automatic rollback slots (a pre-restore snapshot is captured before any non-dry-run restore).
Restore procedure
Section titled “Restore procedure”
From database dumps to an isolated recovery target
Section titled “From database dumps to an isolated recovery target”Do not drop or overwrite the original databases. Prepare new target volumes, private staging storage and a recovery project with no public ingress or device/SMTP egress. Keep originals and encrypted backups intact until acceptance is complete.
- Stop and fence every source writer, including APIs, schedulers, both worker
classes, collectors, plugins, HA peers and external maintenance clients.
Preserve queues, reserved work and Sentinel state before choosing an authority.
Consult
docker/redis/QUEUE-BACKUP-RECOVERY; do not blindly purge or replay work. - Recover the completed encrypted primary/LogDB pair and matching application
volumes. Required files can include report archives, uploaded firmware,
evidence, plugin data and backup artifacts. A database reference to a missing
file is not a restored artifact. Preserve
SECRET_KEYandENCRYPTION_SALTunchanged for this database-level recovery. - Verify GPG decryption completes successfully and validate gzip integrity in
private staging storage before importing. Keep plaintext output private and
encrypted at rest where required. An error at any stage invalidates the target;
do not continue with a partial stream. For shell pipelines, require
pipefailandpsql -v ON_ERROR_STOP=1so intermediate/decode/SQL failures are not hidden. - Restore into compatible empty target databases with general workers stopped.
The scheduled
--no-owner --no-privilegesdumps are not a backup of arbitrary PostgreSQL roles, role passwords, memberships or ACLs. Recover required role configuration separately and verify application access. For a runtime change, followdocker/postgres/UPGRADING; never attach old data volumes directly. - Restore TimescaleDB with the matching extension and schema-qualified
public.timescaledb_pre_restore()/public.timescaledb_post_restore()hooks. The offline runtime migration helper additionally requires background workers disabled on both LogDB helper containers. Restore normal production settings only after the helpers stop. Verify extension objects, jobs, indexes and event payloads, not just table counts. - Apply candidate migrations once while general workers remain stopped. Do not stamp directly to head or start every service to make a migration error go away.
- Follow
docker/TOKEN-REVOCATION-UPGRADING: provision a fresh independentJWT_SIGNING_KEY, different from all previous signing and vault keys, outside the restored snapshot. Apply the same value to all target APIs/workers/schedulers and recreate them under isolation. Preserve vault keys. Revoking session rows is useful defense in depth; incrementing a restored token version by one is not a complete fence against tokens issued after the backup. - Follow
docker/SLA-REPORT-DELIVERY: coordinate a fresh persistentSLA_REPORT_DELIVERY_EPOCH, quarantine old active notices and review source/relay history before explicit schedule reapproval. An unknown mail outcome must not be automatically resent. JWT key rotation does not solve report-mail replay or revoke API keys, OAuth/provider/agent credentials and other non-JWT secrets. - Reconcile customer jobs, authority and device state before starting reviewed workers or opening egress. Verify current queue subscriptions with harmless work. Complete the acceptance checks below, including container recreation, before admitting users. Keep the source stopped and fenced throughout cutover.
The repository’s scripts/check_install_acceptance.py is an automated disposable
Lite rehearsal using synthetic data. See docker/INSTALL-ACCEPTANCE for its exact
arguments and limits. It is not a command that restores a customer installation,
and its passing report does not prove that a site’s backups are complete.
Post-restore validation
Section titled “Post-restore validation”Before reopening traffic, record evidence for the exact target images and backup:
- Both database schemas reach the expected migration state; representative row, event, index and required application-file checks pass.
- Old access, remembered refresh, reset and pending MFA tokens fail, including tokens issued after the restored snapshot. Fresh password login and refresh work, and representative stored credentials decrypt without changing vault keys.
- Restored API keys and other non-JWT credentials have an explicit review; authority or credential revocations lost after the snapshot are reconciled.
- Report archives retain their original bytes and new PDF/CSV reports generate and download correctly. Old text placeholders from v26.09.0 PDF requests must be preserved but refused as PDF downloads; regenerate them after recovery.
- No old/unknown mail or device job is automatically replayed. Only explicitly reviewed schedules and workers resume, under the coordinated current epoch.
- Application readiness and audit-chain checks pass through trusted TLS using the intended account. They do not replace the data and security checks above.
- These checks still pass after recreating the target containers. No original deployment process remains reachable under old authentication settings.
Investigate any audit-chain failure against retained backups; do not rewrite the chain to make validation green. Keeping source data intact is a recovery option, not proof that restarting an old backend or old signing key is safe.
RPO and RTO must be measured
Section titled “RPO and RTO must be measured”The daily schedule is a backup interval, not a guaranteed 24-hour recovery point. Failed or stale backups, delayed off-site copies, missing files and unreconciled queues can increase data loss. Choose the newest verified complete recoverable point and measure its age. Object replication or a green upload alone is not that verification.
No production RTO or RPO is established by a small local rehearsal. Record elapsed time from incident response through key recovery, download, database/file restore, authentication cutover, job reconciliation and successful acceptance. Measure on representative data and failure scenarios, including loss of the original host. Operator-configured WAL streaming needs its own recovery and retention validation; there is no default five-minute guarantee.
Next steps: High Availability describes the single-host failover overlay and its limits; it does not replace cold recovery.
Application vault recovery evidence
Section titled “Application vault recovery evidence”A successful upload is not recovery evidence. Before relying on an application
.fsdnvault, download it through the configured transport and restore it into an
isolated empty database with a different SECRET_KEY and ENCRYPTION_SALT.
Verify login passwords, re-encrypted integration credentials and the device-to-credential
links, as well as controller TLS and polling settings. Reject a corrupted archive and
an incorrect passphrase without creating tenant rows. Record measured recovery time,
archive age and size, and the exact populated contributors; empty sections do not
prove module recovery. This drill complements, but does not replace, restoration of
both PostgreSQL dumps or testing provider-side VM/disk/media recovery.
See remote transport security for the required FTPS certificate trust and SFTP host-key enrollment. An older green connection status does not waive the new transport requirements.
All product names, logos, and brands are property of their respective owners. FreeSDN is an independent project and is not affiliated with or endorsed by the vendors it integrates with. See Trademarks.