High Availability
The Max overlay adds Valkey Sentinel, a PostgreSQL streaming standby and a second API replica. The shipped Compose topology runs on one Docker host. It does not provide protection against host loss or establish enterprise production readiness.
| Component | Mechanism | Operational limit |
|---|---|---|
| Valkey | Two data nodes and three authenticated Sentinels; automatic election | Asynchronous replication can lose acknowledged writes. A partitioned former primary is not fenced. |
| PostgreSQL | Streaming standby | Promotion and application reconnection are manual. Fence the old primary first. |
| API | Two replicas behind an internal nginx proxy and the public Caddy edge | The proxy, edge and Docker host remain shared failure points. |
No load-tested RTO, zero-RPO guarantee or uptime SLA is established by this preset. Recovery observations from small disposable fixtures are evidence for those scenarios only.
Fresh qualification installation
Section titled “Fresh qualification installation”Use the current .env.max.example as the configuration template. Review network overlap and all secrets before creating an isolated installation:
docker compose --env-file .env.max \ -f docker-compose.yml -f docker-compose.ha.yml up -dExisting HA installations must read the shipped docker/redis/HA-UPGRADING first. The database runtime stays on Debian and is adopted in place, so no logical migration is required for PostgreSQL or LogDB. The revised Valkey topology is the breaking part: it deliberately refuses existing AOF/RDB data without its HA state files, and an automated in-place upgrade from the old HA layout is not qualified. docker/postgres/UPGRADING is still the reference if you are returning from an Alpine-based cluster.
Valkey addressing and durable state
Section titled “Valkey addressing and durable state”The default Max network is 172.30.0.0/16. NETWORK_IP_RANGE=172.30.0.0/17 reserves the lower half for dynamic containers. Five fixed addresses outside that pool identify the two data nodes and three Sentinels:
| Setting | Default |
|---|---|
HA_PRIMARY_IP |
172.30.255.10 |
HA_REPLICA_IP |
172.30.255.11 |
HA_SENTINEL_1_IP |
172.30.255.21 |
HA_SENTINEL_2_IP |
172.30.255.22 |
HA_SENTINEL_3_IP |
172.30.255.23 |
Choose a subnet that does not overlap your LAN, VPN or other Docker networks. Change all five addresses, NETWORK_SUBNET, NETWORK_IP_RANGE and FORWARDED_ALLOW_IPS together. Recreating an existing Docker network requires a maintenance window; do not change active node addresses in place.
Numeric addresses prevent Docker’s removal of a stopped service’s DNS record from stalling Sentinel’s resolver. The backend factory and Celery discover the elected primary through those Sentinel addresses. Data nodes and Sentinels retain role and election state on five distinct named volumes, including across container recreation. Keep the volumes and node identities together.
The quorum remains two of three. Loss of the majority must prevent promotion. Both data nodes share the same authenticated configuration and cache/broker memory policy; the current preset does not provide a separate broker cluster. Authentication is required, but the internal Valkey transport is not TLS-encrypted by this preset.
Run isolated fault qualification
Section titled “Run isolated fault qualification”Build the images from the exact candidate under review, then run:
docker build -f docker/Dockerfile.valkey -t freesdn-valkey:qualification dockerdocker build --target production -t freesdn-backend:qualification backendpython3 scripts/check_valkey_ha.py \ --valkey freesdn-valkey:qualification \ --backend freesdn-backend:qualification \ --report valkey-ha.jsonThis creates a unique disposable project with synthetic data, an internal network and no host ports. It tests real master stop, SIGKILL, continued use of existing synchronous and asynchronous application clients, replica rejoin, replacement of data/Sentinel containers, and refusal to promote without a majority. It also checks startup refusal for unmarked data and uncoordinated credential or identity changes. It removes only its own project resources and retains a redacted report and log.
The default qualification subnet is 198.19.240.0/24; select a different unused /24 with --subnet if needed. scripts/redis_failover_drill.sh now forwards these same arguments to the disposable qualification. It no longer pauses containers in an existing installation. Do not use a paused-process test as proof of stopped/recreated-container recovery.
PostgreSQL promotion needs an operator plan
Section titled “PostgreSQL promotion needs an operator plan”A healthy streaming link does not prove safe application cutover. Before promotion, fence the old primary and all writers, verify replication position and acceptable data loss, preserve recoverable backups, and identify every application, pooler and backup-client connection to the old primary. Promote only after those checks, then verify that the new primary is writable and reconnect the reviewed consumers.
Changing DB_HOST alone is insufficient when PgBouncer still targets the old primary. Also, docker compose restart does not reload changed environment variables; reviewed services must be recreated with the new configuration. Preserve old volumes and keep the former primary fenced until a separately rehearsed rejoin or rebuild procedure completes. Never bring up two writable PostgreSQL primaries to resolve a connection problem.
The low-level physical standby qualification (scripts/check_database_ha.py) tests replication, named-volume persistence and manual promotion using synthetic data. It does not implement automatic production database failover.
Before making production HA claims
Section titled “Before making production HA claims”Repeat qualification on the frozen public candidate and on representative infrastructure. Add sustained load, disk pressure, realistic recovery datasets, queue/in-flight-operation accounting, network partitions, loss of a whole host and measured data loss. A multi-host design needs independent failure domains, fencing and an operational recovery plan; moving these Compose containers between hosts is not itself such a design.
Backups and Restore describes cold disaster recovery when failover is insufficient.
All product names, logos, and brands are property of their respective owners. FreeSDN is an independent project and is not affiliated with or endorsed by the vendors it integrates with. See Trademarks.