Skip to content

Backups and Restore

FreeSDN ships backup mechanisms with different scopes. Using the wrong one during an incident is costly - understand each before you need them.

Mechanism What it captures Use case
pg-backup (DB dumps) Application and LogDB database contents, including stored encrypted credentials; not application file volumes, broker state, arbitrary roles/ACLs or external device data Database recovery as one part of a coordinated full-instance recovery
Configuration Backup → .fsdn Config snapshot: Sites, Controllers, Devices, users, automation. No secrets; tied to the source instance key Config portability, version history, dev-to-prod config copy
Configuration Backup → .fsdnvault Supported configuration contributors and their secrets, sealed under an operator passphrase and re-keyed on restore; not every database row, volume, queue or external VM/disk Guided configuration and credential recovery into a compatible empty target

The freesdn-pg-backup container is part of the core stack - no profile required. It dumps both databases daily:

  • freesdn - the primary PostgreSQL DB (19 schemas: access, agents, ai, analytics, audit, backup, cameras, collector, core, devices, enterprise, events, fabric, firewall, gateway, hypervisor, network, voip, vpn)
  • freesdn_logs - the TimescaleDB observability DB

Dumps land in the pg_backups volume with 7-day local retention.

By default, pg-backup fail-closes without GPG configured: it refuses to write unencrypted dumps. This applies to every deployment tier - the guard checks only whether GPG is active, not ENVIRONMENT. Any stack that lacks a valid GPG recipient must set BACKUP_ALLOW_PLAINTEXT=1 explicitly (the homelab .env.lite template includes a commented-out opt-in; uncomment it only if you have no GPG key and understand the risk - pro and max tier templates include GPG placeholder values that you must fill in before running - generate a dedicated keypair, set BACKUP_GPG_RECIPIENT to the real recipient email, and place the exported public key at ./secrets/backup-public.asc - and must not set this flag). Configure a GPG recipient before running in production.

Setup:

  1. On a secure restore host (not the production server), generate a dedicated backup key pair:

    Terminal window
    gpg --full-generate-key
    # Choose RSA 4096, no expiry
    gpg --export --armor backup@example.com > backup-public.asc
  2. Copy backup-public.asc to ./secrets/backup-public.asc in the repo.

  3. Set in your env file:

    Terminal window
    BACKUP_GPG_RECIPIENT=backup@example.com
    BACKUP_GPG_PUBLIC_KEY_PATH=./secrets/backup-public.asc

The container holds only the public key and cannot decrypt its own dumps. Store the private key and its passphrase in a secrets manager completely separate from the production server.

The pg-backup healthcheck requires an enabled backup service, no recorded failure on its latest attempt, and a completed backup pair whose timestamp is within BACKUP_INTERVAL_SECONDS + BACKUP_FRESHNESS_GRACE_SECONDS. The defaults are 24 hours plus 2 hours. Both the application and LogDB final files must exist and be nonempty. A future, malformed or mismatched completion timestamp fails health. Old files or a recently created backup directory do not establish health.

The service atomically replaces its local status file at /backups/.metrics/pg_backup.prom. A failed dump retains the previous success time but sets freesdn_pg_backup_last_attempt_failed to 1. Disabled encryption configuration sets freesdn_pg_backup_enabled to 0 and is unhealthy; the service still idles without writing plaintext. The completion pointer at /backups/.metrics/pg_backup.last_success identifies the pair covered by that success time. A completed dump is not evidence that a restore will succeed.

Enable the metrics profile to run the internal backup-metrics exporter and Prometheus scrape job freesdn-backups. The exporter reads both volumes without write permission, runs without root or Linux capabilities, and publishes no host port. Only /metrics and /healthz are served; archive contents, paths and credentials are not exposed. /healthz means the exporter is reachable, not that backups are healthy.

The shipped alerts cover missing exporter observations, malformed status, missing archive pairs, disabled backups, failed attempts, stale completion, explicitly unencrypted backups, and required off-site status. Freshness thresholds follow the configured intervals. Missing status produces explicit zero-valued validity metrics rather than disappearing time series. Alerts generally require 5 minutes of continuous failure; stale/retention alerts use 15 minutes and the plaintext warning uses 30 minutes. First runs that have not completed yet are reported as unavailable; large initial dumps can therefore alert before finishing.

Alert delivery is not configured automatically. Prometheus evaluates these rules, but an operator must configure Alertmanager and a notification receiver to receive pages or email. Test that delivery path before relying on it.

The local completion timestamp is not a worker heartbeat. If a worker stops after a successful cycle, age-based alerts fire when the configured limit is exceeded. A completed dump or verified remote copy still does not prove that a restore will succeed.

The configured service begins a backup cycle when it starts. With a verified GPG recipient and public key, restart only pg-backup and wait for the new completed pair before proceeding:

Terminal window
docker compose --env-file .env.pro restart pg-backup
docker compose --env-file .env.pro ps pg-backup
docker compose --env-file .env.pro logs --tail=50 pg-backup

Use the selected installation’s actual environment file and project. A zero exit from restart does not mean the backup finished. Require fresh completion metadata and nonempty files for both databases, verify encryption and retained keys, then rehearse restoring the pair. Treat raw logs and file paths as private.

Do not run a separate pg_dump | gzip shell to bypass the backup entrypoint: that path does not inherit the entrypoint’s encryption checks and can create plaintext dumps in a production backup volume. A consistent recovery point across the databases, files and broker also requires stopping/accounting for writers; two separately completed daily dumps do not establish that global consistency.

List files and read the completion pointer privately to select an actual completed pair. Do not infer the database filename from a hardcoded example date. Copy files as binary data, without allocating a TTY:

Terminal window
docker compose --env-file .env.pro exec -T pg-backup ls -lh /backups
docker compose --env-file .env.pro exec -T pg-backup cat /backups/.metrics/pg_backup.last_success

Use docker compose cp with each reviewed exact filename into a protected local directory, or a non-TTY binary stream with exec -T. Verify the copied bytes against the source and retain the pair, matching vault configuration and required application files together. Never publish dumps, secret configuration or private decryption keys in an issue or release artifact. Test recovery using a separately retained private key; keeping only the public encryption key is insufficient.

Enable dr and metrics in COMPOSE_PROFILES, and configure a dedicated rclone remote. See Compose Profiles for setup steps. Set BACKUP_OFFSITE_REQUIRED=1 to require remote completion evidence even before the first run, especially if using --profile dr on the command line. The exporter also infers this requirement from COMPOSE_PROFILES=...,dr or an existing remote status file. Removing dr after it has been used does not silently clear its monitoring requirement.

The off-site service copies retained, nonempty database pairs with generated filenames. It excludes .partial files, orphaned dumps, and manually named files. It snapshots the local completion pointer and requires both files of that pair. After copying, it downloads remote contents for comparison and checks that every selected source file appears in the verified match list. Empty or vanished source files cannot produce a false completion. Once verified, it writes and reads back _completed/<batch>.<extension>.txt on the remote, then atomically records status in the separate pg_backup_offsite_status volume. Archives remain mounted read-only.

Verification reads all selected retained backup data on each cycle. Budget for network bandwidth, provider requests, egress charges and completion time. A failed copy, content check or marker operation preserves the previous success and records a failed latest attempt. Rechecking the same old backup updates the verification time but does not reset its source age. These observations apply to the last verification; remote loss after that check is detected on a later run.

Off-site retention defaults to 30 days, independently of local retention. The currently acknowledged pair and marker are protected from that sweep, even when old. Retention errors are reported separately and fail container health without erasing successful copy evidence. The rclone --immutable option prevents this client from replacing different existing content; it is not storage-enforced immutability. Configure object lock/versioning, separate credentials and an independent retention policy with your storage provider. Do not weaken object-lock protection merely to silence a deletion warning.

Recovery points depend on both schedules. Daily local dumps can leave about 24 hours of data since the previous snapshot. A separate daily off-site copy loop can add nearly another 24 hours, plus dump/transfer/verification time. Neither interval is a guaranteed RPO or RTO. With defaults, remote completion-age alerts use 24h local + 24h copy + 2h grace. Use shorter intervals or a separately tested continuous backup design when the recovery objective is tighter. Measure it with restores from the actual remote storage and keep the decryption keys off the production server.

The Configuration Backup module (at /backup in the UI) produces portable .fsdn archives. Use them to:

  • Migrate configuration from a staging instance to production
  • Snapshot configuration before a major change
  • Seed a new Site with an existing Site’s configuration

A .fsdn archive carries no credentials. After restoring a config archive to a new instance, re-enter all module secrets (Controller passwords, NVR credentials, etc.).

To move supported configuration and its secrets, the module produces a .fsdnvault archive, sealed under an operator passphrase and re-keyed using the destination’s encryption configuration on restore. It also drives the setup wizard’s “Restore from a backup instead” branch on a compatible fresh instance. Restored organization/admin configuration must pass login, credential, contributor and integration checks. This archive does not restore external VM disks, media, broker state or every database row and application file. See Configuration Backup → secure full backup for the create/restore flow and credential-handling requirements.

The module supports selective restore (individual sections), strict semver schema gating (a mis-versioned section is rejected, not silently applied), and automatic rollback slots (a pre-restore snapshot is captured before any non-dry-run restore).

From database dumps to an isolated recovery target

Section titled “From database dumps to an isolated recovery target”

Do not drop or overwrite the original databases. Prepare new target volumes, private staging storage and a recovery project with no public ingress or device/SMTP egress. Keep originals and encrypted backups intact until acceptance is complete.

  1. Stop and fence every source writer, including APIs, schedulers, both worker classes, collectors, plugins, HA peers and external maintenance clients. Preserve queues, reserved work and Sentinel state before choosing an authority. Consult docker/redis/QUEUE-BACKUP-RECOVERY; do not blindly purge or replay work.
  2. Recover the completed encrypted primary/LogDB pair and matching application volumes. Required files can include report archives, uploaded firmware, evidence, plugin data and backup artifacts. A database reference to a missing file is not a restored artifact. Preserve SECRET_KEY and ENCRYPTION_SALT unchanged for this database-level recovery.
  3. Verify GPG decryption completes successfully and validate gzip integrity in private staging storage before importing. Keep plaintext output private and encrypted at rest where required. An error at any stage invalidates the target; do not continue with a partial stream. For shell pipelines, require pipefail and psql -v ON_ERROR_STOP=1 so intermediate/decode/SQL failures are not hidden.
  4. Restore into compatible empty target databases with general workers stopped. The scheduled --no-owner --no-privileges dumps are not a backup of arbitrary PostgreSQL roles, role passwords, memberships or ACLs. Recover required role configuration separately and verify application access. For a runtime change, follow docker/postgres/UPGRADING; never attach old data volumes directly.
  5. Restore TimescaleDB with the matching extension and schema-qualified public.timescaledb_pre_restore() / public.timescaledb_post_restore() hooks. The offline runtime migration helper additionally requires background workers disabled on both LogDB helper containers. Restore normal production settings only after the helpers stop. Verify extension objects, jobs, indexes and event payloads, not just table counts.
  6. Apply candidate migrations once while general workers remain stopped. Do not stamp directly to head or start every service to make a migration error go away.
  7. Follow docker/TOKEN-REVOCATION-UPGRADING: provision a fresh independent JWT_SIGNING_KEY, different from all previous signing and vault keys, outside the restored snapshot. Apply the same value to all target APIs/workers/schedulers and recreate them under isolation. Preserve vault keys. Revoking session rows is useful defense in depth; incrementing a restored token version by one is not a complete fence against tokens issued after the backup.
  8. Follow docker/SLA-REPORT-DELIVERY: coordinate a fresh persistent SLA_REPORT_DELIVERY_EPOCH, quarantine old active notices and review source/relay history before explicit schedule reapproval. An unknown mail outcome must not be automatically resent. JWT key rotation does not solve report-mail replay or revoke API keys, OAuth/provider/agent credentials and other non-JWT secrets.
  9. Reconcile customer jobs, authority and device state before starting reviewed workers or opening egress. Verify current queue subscriptions with harmless work. Complete the acceptance checks below, including container recreation, before admitting users. Keep the source stopped and fenced throughout cutover.

The repository’s scripts/check_install_acceptance.py is an automated disposable Lite rehearsal using synthetic data. See docker/INSTALL-ACCEPTANCE for its exact arguments and limits. It is not a command that restores a customer installation, and its passing report does not prove that a site’s backups are complete.

Before reopening traffic, record evidence for the exact target images and backup:

  • Both database schemas reach the expected migration state; representative row, event, index and required application-file checks pass.
  • Old access, remembered refresh, reset and pending MFA tokens fail, including tokens issued after the restored snapshot. Fresh password login and refresh work, and representative stored credentials decrypt without changing vault keys.
  • Restored API keys and other non-JWT credentials have an explicit review; authority or credential revocations lost after the snapshot are reconciled.
  • Report archives retain their original bytes and new PDF/CSV reports generate and download correctly. Old text placeholders from v26.09.0 PDF requests must be preserved but refused as PDF downloads; regenerate them after recovery.
  • No old/unknown mail or device job is automatically replayed. Only explicitly reviewed schedules and workers resume, under the coordinated current epoch.
  • Application readiness and audit-chain checks pass through trusted TLS using the intended account. They do not replace the data and security checks above.
  • These checks still pass after recreating the target containers. No original deployment process remains reachable under old authentication settings.

Investigate any audit-chain failure against retained backups; do not rewrite the chain to make validation green. Keeping source data intact is a recovery option, not proof that restarting an old backend or old signing key is safe.

The daily schedule is a backup interval, not a guaranteed 24-hour recovery point. Failed or stale backups, delayed off-site copies, missing files and unreconciled queues can increase data loss. Choose the newest verified complete recoverable point and measure its age. Object replication or a green upload alone is not that verification.

No production RTO or RPO is established by a small local rehearsal. Record elapsed time from incident response through key recovery, download, database/file restore, authentication cutover, job reconciliation and successful acceptance. Measure on representative data and failure scenarios, including loss of the original host. Operator-configured WAL streaming needs its own recovery and retention validation; there is no default five-minute guarantee.

Next steps: High Availability describes the single-host failover overlay and its limits; it does not replace cold recovery.

A successful upload is not recovery evidence. Before relying on an application .fsdnvault, download it through the configured transport and restore it into an isolated empty database with a different SECRET_KEY and ENCRYPTION_SALT. Verify login passwords, re-encrypted integration credentials and the device-to-credential links, as well as controller TLS and polling settings. Reject a corrupted archive and an incorrect passphrase without creating tenant rows. Record measured recovery time, archive age and size, and the exact populated contributors; empty sections do not prove module recovery. This drill complements, but does not replace, restoration of both PostgreSQL dumps or testing provider-side VM/disk/media recovery.

See remote transport security for the required FTPS certificate trust and SFTP host-key enrollment. An older green connection status does not waive the new transport requirements.

All product names, logos, and brands are property of their respective owners. FreeSDN is an independent project and is not affiliated with or endorsed by the vendors it integrates with. See Trademarks.