Unreleased
These notes describe the changes since v1.1.0 that have not yet been released. The Changelog summarizes them.
Overview
This release adds recovery of an existing cluster from its pgBackRest repository, and reworks when the collection takes a backup so that redeploying a cluster can never discard the recovery point its repository was holding.
Its main features are:
- A recovery playbook that rebuilds an HA cluster from its pgBackRest
repositories, optionally to a point in time: the Ultra-HA deployment with a
new
recover_postgresrole in place ofsetup_postgres, preceded by a newwipe_clusterrole and followed, once the result is right, by a commit. See Recovering a Cluster from Backup. - A new
finalize_backrestrole that creates the stanza, the first backup and the backup schedule at the end of a deployment, and takes that first backup only when the repository holds no backup that can restore the cluster running now. - An etcd certificate authority that can be supplied from Ansible Vault, so a controller other than the one that first deployed the cluster can manage it.
backup_repo_cipherno longer has a default, because the default could be derived by anyone who knew the cluster's name.
Upgrading from v1.1.0
A playbook written for v1.1.0 needs changes to work with this release. Read this
section before running the new collection against an existing cluster, or
before deploying with a playbook you wrote for v1.1.0. The sample playbooks
under sample-playbooks/ already include every change below.
Reorder the backup roles
setup_backrest now writes configuration files only, and the steps that need a
running cluster moved to the new finalize_backrest role. A playbook that still
ends with setup_backrest and never applies finalize_backrest runs without
an error and sets up no backups. Nothing creates the stanza, takes the first
backup, or installs the cron entries, and a non-HA cluster never gets its
archive_command.
An HA cluster fills its disk
On an HA cluster the result is worse than having no backups. Patroni now
renders the pgBackRest archive_command into the Postgres configuration
itself, so Postgres starts archiving as soon as Patroni starts it. Without a
stanza every archive attempt fails, Postgres keeps every WAL segment until
one succeeds, and pg_wal grows until the disk is full.
Archive failures before the stanza exists are expected
Even with the roles in the right order, the Postgres log on an HA cluster
shows failed archive-push attempts from the moment Patroni starts Postgres
until finalize_backrest creates the stanza at the end of the deployment.
This is expected. Postgres keeps the WAL it could not archive, and the
archiver sends all of it to the repository once the stanza exists. If the
failures continue after finalize_backrest has run, look into them: they
are no longer part of the deployment.
Make the following changes to every playbook that configures backups:
- In the play for the
pgedgehosts, movesetup_backrestso it comes beforesetup_postgres. On an HA cluster that also puts it beforesetup_patroni. - Keep
setup_backrestin the play for thebackuphosts. - Add a final play that applies
finalize_backrestto both groups, aftersetup_pgedgeand after the backup servers are configured.
The roles in an Ultra-HA playbook then run in the following order. The
collections keys and the when conditions on the etcd and pgBouncer roles are
left out here; sample-playbooks/ultra-ha/playbook.yaml has them.
- hosts: all
roles:
- init_server
- hosts: pgedge
roles:
- install_repos
- install_pgedge
- install_etcd
- install_patroni
- install_backrest
- setup_backrest # moved: before setup_postgres
- setup_postgres
- setup_etcd
- setup_patroni
- hosts: haproxy
roles:
- setup_haproxy
- hosts: pgedge
roles:
- setup_pgedge
- hosts: backup
roles:
- install_repos
- install_backrest
- setup_backrest
- hosts: pgedge:backup # new: the last play in the playbook
roles:
- finalize_backrest
Running finalize_backrest against an existing cluster is safe. It takes a full
backup only when the repository holds no backup that can restore the cluster
running now, so a repository that already has backups of it keeps them. It
also checks the archive command in Patroni's configuration store, and changes
it only when the store does not already point at pgBackRest.
full_backup_schedule and diff_backup_schedule now belong to
finalize_backrest, which installs the cron entries. Values set in the
inventory still apply. A playbook that passes them as role variables to
setup_backrest must pass them to finalize_backrest instead, or the
schedule falls back to the defaults.
Set backup_repo_cipher
backup_repo_cipher no longer has a default. When a repository is configured
and backup_repo_cipher_type is aes-256-cbc, which is the default,
init_server stops the playbook until the parameter is set. This applies to
existing clusters as well. They keep running, but the next playbook run
against one stops in init_server.
The repository of an existing cluster is encrypted with the old derived
password. You can recover that password, because anyone could derive it.
Upgrading a cluster deployed before this was required
gives the command. Store the value in Ansible Vault and set
backup_repo_cipher from it.
The derived password included the zone, so every zone of a cluster that used
the default has a different password. The sample inventories and the role
documentation set one backup_repo_cipher for the whole cluster. That is fine
for a new cluster, because nothing requires the zones to differ. It is wrong for
an upgraded multi-zone cluster. A single value matches at most one zone's
repository. The playbook rewrites pgbackrest.conf on every other zone with the
wrong password. setup_backrest now notices when the new file cannot read a
stanza that the current file can, and stops before writing it, so archiving on
those zones' primaries keeps working, but the playbook cannot go further until
each zone has its own password.
Recover the password of each zone, store each one in Ansible Vault, and select
it by zone. Keep the setting under all so that the backup servers resolve the
same value as the nodes they serve:
# vault.yaml
vault_backup_repo_ciphers:
1: <recovered password for zone 1>
2: <recovered password for zone 2>
all:
vars:
backup_repo_cipher: "{{ vault_backup_repo_ciphers[zone | int] }}"
A cluster that already set backup_repo_cipher explicitly does not need this
step, because its repositories already use the value in its inventory.
Treat a recovered password as compromised. It lets you keep reading the existing repository, but you should plan to re-encrypt the repository.
A new cluster needs a password of your own choosing. If the storage layer
encrypts the repository, for example with S3 default bucket encryption, you
can set backup_repo_cipher_type: none instead.
Supply the etcd certificate authority to other controllers
Every play that signs etcd or Patroni certificates now reads the certificate
authority that the running etcd members trust, and stops unless the controller
holds the same one. A cluster deployed from one controller and managed from
another, such as a second workstation or a CI runner, needs etcd_ca_cert and
etcd_ca_key set from Ansible Vault before that controller can run the
playbook. Previously a controller without the authority generated a new one and
reissued the cluster's certificates against it, and Patroni then could not
reach etcd. The play also stops when the controller's tls/etcd/ holds the
certificate without its private key, or with a key that does not match it,
which previously generated a new authority over the staged one.
etcd_ca_cert and etcd_ca_key
describes how to capture the authority from the original controller.
A cluster that is only ever managed from the controller that deployed it needs no change.
Any cluster whose Patroni uses etcd can still meet one new stop. The check asks
every host in the pgedge group, not only the hosts the play runs on, so
setup_patroni, and setup_etcd whenever it builds a node, now stop when any
pgEdge host in the inventory is unreachable, naming that host, even under
--limit. Previously an unreachable host dropped out of the play and the rest
carried on. Bring the host back, or remove it from the inventory if it has left
the cluster for good.
Gather facts for every host first
Roles read other hosts' addresses from the facts gathered for those hosts, so
the first play in a playbook must target every host in the inventory, without
become. Gathering facts without become also records the login account
rather than root. The Ultra-HA sample already applies init_server to all
in its first play. The simple-cluster sample now starts with a play that only
pings every host. If your playbook's first play targets a single group, add the
same play to the top:
- hosts: all
any_errors_fatal: true
tasks:
- name: Confirm every host in the inventory is reachable
ping:
Changes
Added
- New
wipe_clusterrole andsample-playbooks/wipe-cluster/playbook tear a cluster down for a redeployment or a recovery. They stop Postgres, Patroni, pgBouncer and the collection's etcd, remove the cluster from any other configuration store withpatronictl, remove the backup schedule, and erase the components' configuration and data. No repository is touched, and the wipe is refused unlesswipe_confirmis set and some zone's repository holds a backup of the cluster in it.wipe_without_backuplifts the second condition for a cluster whose data is not worth keeping. Every pgEdge node is asked for a backup, and a zone is judged by whichever of its nodes holds its cluster, so a first node that lost its disk cannot make a zone that failed over look empty. An external configuration store the wipe cannot read stops it before anything is erased. The zone's cluster is removed from that store even when the store lists no members, since members expire within Patroni'sttlbut the keys recording the cluster as initialized do not, andpatronictlmust then report the cluster asuninitialized. A node with no Patroni configuration, such as a freshly provisioned host, reaches the store through a temporary one built from the inventory'spatroni_dcssettings. (EE-39) - New
recover_postgresrole builds the cluster's Postgres from a backup, in place ofsetup_postgres. It appliessetup_postgresto every node, with the newbackup_repo_reuseset so that the repository identity check lets it build a cluster beside the existing stanza, restoresrecovery_nodefrom its zone's repository and promotes it outside Patroni, and strips the restored node's stale Spock metadata. One zone is restored and every other zone copies it across Spock, because each zone's stanza restores to its own moment, and zones restored separately have no common position to replicate forward from. It then runspgbackrest stanza-upgradeon every zone whose stanza does not describe the cluster now in it. A rebuilt zone comes back with a system identifier its stanza has never seen, so PgBackRest would refuse to archive there, andpg_walwould keep every segment through the Spock copy. The upgrade only adds an entry to the stanza's history: every earlier backup stays in place and restorable, which is what lets a recovery be retried from that zone before it is committed. Becausestanza-upgraderuns on the repository host, the role refuses, before building anything, a rebuilt zone whose SSH repository is named only bybackup_host; adding that server to thebackupgroup lets the recovery upgrade it. (EE-39) - New
sample-playbooks/ultra-ha-recover/holds the recovery playbook, which is the Ultra-HA deployment playbook withrecover_postgresin place ofsetup_postgres, andcommit-restore.yaml, which appliesfinalize_backrestonce the recovered cluster is the one to keep. The recovery takes no backups and leaves the schedule off, so it can be repeated to another target or from another zone with every backup still in place. (EE-39) - New
recovery_node,recovery_target_type,recovery_target,recovery_target_timelineandrecovery_backup_setparameters drive a recovery.recover_postgresreads them, and thesample-playbooks/ultra-ha-recover/playbook also readsrecovery_nodeto choose the zone the others are seeded from. (EE-39) - New
pgedge_seed_zoneparameter forsetup_pgedgenames a zone that already holds the cluster's data. Every other zone copies it withsynchronize_structureandsynchronize_databefore the rest of the mesh is built. The recovery playbook sets it to the restored zone. (EE-39) - New
pgedge_seed_stall_minutesparameter (default 15) forsetup_pgedgebounds the wait for that copy by progress, not time. The wait pollsspock.sub_show_status()and the sync status instead of callingspock.sub_wait_for_sync(), which keeps waiting while Spock retries a failed copy. It fails at once on a disabled subscription, after two minutes of an apply worker that will not stay up (a bad DSN or password, or a copy that broke partway), or after the stall window without progress. It reports the subscription status and the Spock lines from the Postgres log.pgedge_seed_max_hours(default 24) remains the backstop. (EE-39) - New
patroni_replica_from_backupparameter makes Patroni rebuild a replica with a pgBackRest delta restore instead of a freshpg_basebackupfrom its zone's leader, falling back topg_basebackup(pg_cloneclusteron Debian) if the restore fails. Off by default. The restore runs through a scriptsetup_patroniinstalls at/usr/local/bin/patroni_pgbackrest_replica, which first asks the leader for its system identifier and timeline history and fails, without restoring, unless the newest backup is of the leader's cluster and ends on its history. After a point-in-time recovery that stopped before the newest backup, the replicas therefore fall back instead of restoring a backup they cannot follow. On Debian the script also creates the/etc/postgresqlconfiguration directory after the restore. (EE-39) - New
finalize_backrestrole initializes the backup repository and the backup schedule: the backup database user, the stanza, the first backup and the cron entries. It is applied at the end of a deployment, where a running cluster exists for it to act on. The backup user, and in S3 mode the stanza and the first backup, come from each zone's current Patroni leader, so the role works on an HA cluster that has failed over away from its first node. (EE-39) role_configgained arepo_identitytask file that compares a node's cluster against the one its stanza describes, and refuses when they disagree.setup_postgresasks it before initializing a data directory andfinalize_backrestasks it before writing to the repository. The early check turns the disaster-recovery case — replacement hardware, the original inventory, and a repository that outlived the cluster — into a stop with an explanation rather than an empty cluster that can never archive to its own repository. It skips a repository it cannot reach, since an SSH repository is not reachable from a pgEdge node that early; the late check, which gates a decision to write, treats an unreadable repository as fatal. (EE-39)finalize_backresttakes a full backup when the stanza's newest backup of the cluster cannot restore it: after a point-in-time recovery the newest backups can lie on the timeline the recovery abandoned, past the point the cluster branched away. It reads the cluster's timeline history from the zone's primary (from its first pgEdge node when a backup server asks). A failover branches after every backup, so it never causes one. (EE-39)finalize_backrestcounts only backups of the cluster running now, matched by system identifier, so a stanza that holds only backups of a cluster a recovery replaced is treated as having none. (EE-39)- The Ultra-HA end-to-end test now verifies the backup surface rather than
printing it. It asserts that the running server's
archive_commandinvokes pgBackRest rather than the template's/bin/trueplaceholder, that WAL archiving is currently succeeding according topg_stat_archiver, and that the repository holds exactly one healthy stanza with exactly one full backup. It then re-appliesfinalize_backrestand asserts the repository holds the same backups afterwards, which is the assertion that would have caught the destructive behaviour this release removes: a second full backup expires its predecessor under the default retention, so a repository rewritten that way still holds one backup and only its label changes. The Ultra-HA workflow'sbackupinput now runs this against an S3 repository on MinIO as well as an SSH one. (EE-39) - The recovery waits for the restored node's replay by watching whether it is
still making progress rather than by counting attempts. How long a replay
takes is a property of the database, so any fixed budget is wrong for some
cluster. The wait follows the node's log and
pg_control's timestamp, and gives up only when neither has moved forrecovery_stall_minutes(default 15), or at once if Postgres stops.recovery_max_hours(default 24) is a hard ceiling. The restore and the wait run underasync, so neither is lost with an SSH connection, and are polled everyrecovery_poll_seconds(default 30). (EE-39) - New
etcd_ca_certandetcd_ca_keysupply the certificate authority that signs etcd's certificates and every node's Patroni client certificate from Ansible Vault. Previously the authority existed only on the controller that first deployed the cluster, so any other controller — including a fresh CI runner — could not add a replica, rebuild a node, recover the cluster, or re-run the deployment. The etcd configuration page now explains how to generate a new authority or capture an existing one, and how to vault it. Both parameters are empty by default and an inventory that sets neither behaves exactly as before. A supplied authority that differs from the one already staged is refused rather than applied, because signing against a different authority than the running etcd trusts leaves no node able to reach the store. (EE-39) - New
tests/render/check-patroni.pyrenders the Patroni template offline across every combination of the backup switches, and asserts that a template rendered without facts for the proxy and backup hosts fails rather than emitting HBA rules with no addresses in them. Run by both workflows and both local harnesses. (EE-39) - New
tests/render/check-pgbackrest.pyrenderspgbackrest.confoffline for both cipher types and for SSH and S3 repositories, including the S3-compatible settings. Run by both workflows and both local harnesses. (EE-39) - New
tests/run-recovery-test.sh,tests/playbooks/seed-recovery.yml,tests/verify/verify-recovery.ymlandtests/verify/verify-commit.ymlexercise a recovery against a cluster the end-to-end harness has already deployed: wipe, recover, verify, commit, verify. The seed writes rows after the deployment's backup, so they exist only in archived WAL: a recovery has to replay the archive to return them to the restored zone and carry them across Spock to the zones rebuilt empty. The seed also records each zone's system identifier and timeline on the controller, and the verification asserts the restored zone kept its identifier and moved to a later timeline while every rebuilt zone has a new one. A second pass,lsn, seeds a further batch after a recorded WAL position, recovers to that position and asserts the batch stayed out; the Ultra-HA workflow'srecover_toinput pickslatest,lsnorboth. Every zone-wide check runs against the leader Patroni names, and fails a zone with no leader rather than skipping it. The verification also asserts the replication mesh was rebuilt to exactly the expected size, that every replica came back, that writes made afterwards reach every zone, that each zone can archive again, and that the recovery took no backup. After the commit it asserts every zone has a backup of the cluster it runs. (EE-39) - New
backup_stanzaandbackup_repo_configuredvariables inrole_config, so that the roles which now share them cannot drift apart. (EE-39) - New
uri_style,storage_ca_file,storage_portandstorage_verify_tlskeys inbackup_repo_paramsset PgBackRest'srepo1-s3-uri-style,repo1-storage-ca-file,repo1-storage-portandrepo1-storage-verify-tls, which an S3-compatible store such as MinIO usually needs: path-style addressing, a port of its own, and a certificate authority for an endpoint whose certificate is privately signed -- or, in a test environment, no certificate check at all. All four are empty by default and omitted from the configuration when empty, so an AWS S3 repository renders exactly as before.backup_repo_paramsand the mergedbackup_paramsmoved fromsetup_backresttorole_config, becauseinit_servernow validates them. (EE-39) init_serverrefuses an S3 repository whosebackup_repo_paramsleaves the credentials, bucket, region or endpoint empty, or gives auri_styleother thanhostorpath, astorage_portoutside 1 to 65535, or astorage_verify_tlsthat is not a boolean. Empty credentials used to surface only whenfinalize_backrestfirst wrote to the repository, at the very end of the deployment. (EE-39)init_serverrefuses a zone that uses an S3 repository and also has a host in thebackupgroup. A backup server only serves an SSH repository, and in S3 mode no SSH keys are exchanged, so the server failed partway through the deployment trying to reach nodes it was never given access to. (EE-39)
Changed
setup_backrestno longer replaces apgbackrest.confthat can read the stanza with one that cannot. A wrongbackup_repo_cipheror object-store key used to be written straight over a working file, which broke archiving on the live primaries at once while the run carried on throughsetup_patroniandsetup_pgedge. The new file is now rendered beside the old one and both are asked to read the stanza first. The previous file is also kept as a backup each time it changes, with the same owner and mode. (EE-39)- The initial backup is now taken only when the repository holds no backup that
can restore the cluster running now, rather than on every run.
full_backup_countdefaults to1, so a full backup taken against a repository that already had one expired the previous full backup and its WAL the moment it completed — discarding the recovery point a redeployed cluster was about to be restored from. The decision is now made from the repository's contents rather than from where the role sits in a playbook. (EE-39) setup_backrestnow writes files only and touches Postgres not at all, and moves beforesetup_postgresin the role order. An HA cluster gets itsarchive_commandfrom the Patroni configuration, so Postgres starts archiving the moment Patroni starts it, and apgbackrest.confwritten later left a window in which every archive attempt failed. Placing it beforesetup_postgresalso lets that role ask the repository whether it already holds a cluster before initializing a data directory. A non-HA cluster'sarchive_commandandrestore_commandmoved tofinalize_backrest, which runs when Postgres is up to be reloaded. (EE-39)setup_postgreson Debian now removes the configuration files of a cluster whose configuration directory outlived its data directory before creating it again, sincepg_createclusterrefused to recreate a cluster whose configuration was still in/etc/postgresql. This includes a cluster namedmain, which was previously never created by the role and so was left without a cluster once its data directory was erased. It also letscluster_name: mainwith apg_dataof its own replace the package'smaincluster on a fresh install: the role stops that cluster and replaces its configuration, and leaves its data directory where it is. (EE-39)archive_commandandrestore_commandfor an HA cluster now come fromsetup_patroni's configuration template rather than being patched into the Patroni configuration store bysetup_backrest. A value held only in the store is lost when the store is rebuilt, which is what a recovery does: a recovered cluster came back witharchive_commandreverted to/bin/trueand silently stopped archiving. The patch insetup_backrest'sconfig_postgres_ha.yamlhas been replaced by a check infinalize_backrest. Patroni applies the template only when it bootstraps a cluster, so an HA cluster deployed without a backup server and given one later still has/bin/truein its configuration store.finalize_backrestnow reads the store, patches it only when it disagrees, and waits for Postgres to take up the new command before the first backup. (EE-39)- The Patroni template now spells replica creation as
create_replica_methodsrather than the legacycreate_replica_method. Patroni reads both, although its configuration validator knows only the new spelling. Debian replicas are still built withpg_cloneclusterby default, as before. Debian needs it: it creates the/etc/postgresqlconfiguration directory thatbasebackupleaves out. Withpatroni_replica_from_backupenabled, a Debian replica tries the pgBackRest restore first and falls back topg_clonecluster. (EE-39)
Security
backup_repo_cipherno longer defaults to a value derived fromcluster_nameandzone. Both are public, so the password protecting every backup was reproducible by anyone who knew the cluster's name — encryption in form only. It cannot be defaulted at all: a generated password must be identical on every run, or the repository stops being readable, and must also be unguessable, and nothing can be both.init_servernow requires one, the way it already requires the other passwords, and the parameter moved torole_configso that validation and the configuration template read the same value.
Existing clusters keep running, but the next playbook run against one stops
in init_server until the parameter is set. Their repositories are encrypted
with the old derived password, and the derivation was public, so it can be
recovered — Backup Configuration
gives the command. Treat a recovered value as compromised and plan to
re-encrypt. (EE-39)
- A collection tarball built with make build no longer includes private keys
left in the tree by local runs. ansible-galaxy ignores .gitignore, so the
etcd CA and node keys the sample playbooks write under tls/, the SSH host
keys under host-keys/, and the test harness's SSH key were all packaged.
galaxy.template.yml, from which make build generates galaxy.yml, now
excludes them, along with all of tests/ and other local files. Anyone who
built and shared a tarball from a tree where these existed should treat the
keys in it as exposed. (EE-39)
- Running a playbook with --diff no longer prints secrets. The tasks that
write patroni.yml, pgbackrest.conf and the .pgpass entries showed their
changes in the diff output, and those carry the database passwords, the
repository cipher and any object-store secret key. (EE-39)
Fixed
backup_repo_cipher_type: nonenow produces a configuration PgBackRest accepts. The repository template emittedrepo1-cipher-passunconditionally, and PgBackRest rejects a password alongside a cipher type ofnone, so leaving encryption to the storage layer was not expressible — which is the arrangement that permits key rotation, since PgBackRest cannot rotate its own cipher. The password line is now emitted only where something is encrypting, andinit_serverrejects a password set where nothing is. (EE-39)backup_repo_userandbackup_repo_pathno longer derive fromansible_user_id. That is a fact, and a fact records which account the setup module ran as, so a play that gathers facts withbecomerecordsroot-- and because fact gathering is smart by default, one such play poisons the value for every later play in the same run. The repository owner becamerootand its path/home/root. PgBackRest refuses to run as root, which is the only reason this surfaced at all rather than quietly building a repository somewhere nobody would look for it. Both now derive from a newconnection_userinrole_config, which prefers theansible_userconnection variable and falls back to the fact only where no login user is configured.init_servercompares against the same value. (EE-39)setup_backrestno longer skips its client configuration entirely whenbackup_repo_typeiss3. The role gated client setup on a backup server being named, and an S3 repository names none, so S3 clusters were left with nopgbackrest.conf, no archive command and no backups, with nothing reporting it. The gate is now whether a repository is configured at all. (EE-39)- Adding or rebuilding a node from a controller without the cluster's certificate authority no longer generates a new one. The etcd setup skipped nodes that already ran etcd but not the node being added, so it minted an authority for that node, and every node's Patroni client certificate was then reissued against an authority the running etcd did not trust. Every play that signs certificates now reads the authority each existing etcd member trusts and stops unless the controller holds that one. (EE-39)
setup_patronifinds the primary when Postgres listens on a port other than 5432.patronictlthen shows each member's host ashost:port, and the wait for the primary compared that with the bare inventory name, so it never saw the primary come up and failed after its retries. The port is now stripped before the comparison. (EE-39)make buildrebuilds the tarball whenever a shipped file changes or is deleted, including role templates and scripts, the top-levelmeta/, doc pages and sample playbooks. It compared only the roles' YAML files that still existed, so an edit to anything else, or a deletion, left a stale tarball in place — under the same name, since a dirty tree keeps its version string — andmake installreinstalled the old content. (EE-39)