finalize_backrest
The finalize_backrest role initializes the PgBackRest repository and the backup
schedule for a cluster that is up and running. It creates the backup database
user, creates the repository stanza if it does not exist, takes a full backup
only when the repository holds no backup that can restore the cluster running
now, and installs the cron entries for scheduled full and differential backups.
It is the second half of backup setup. setup_backrest writes configuration and
needs nothing running, so it is applied before Patroni starts Postgres;
everything this role does needs a live cluster to talk to, so it is applied last.
On an HA cluster Postgres starts archiving before this role runs, and every
archive-push fails until the stanza exists; the failures stop, and the WAL
Postgres kept is archived, once this role has created the stanza.
The role performs the following tasks on inventory hosts:
- Compare the cluster's system identifier against the identifiers the stanza has described, and stop if none matches.
- On an HA cluster, find each zone's Patroni leader, and stop if there is not exactly one among the zone's nodes.
- Create the
backup_userPostgreSQL role withpg_checkpointprivileges, and for non-HA clusters add the matchingpg_hba.confentries. - Create the repository stanza when the repository does not already have one.
- For non-HA clusters, set the Postgres
archive_commandandrestore_commandto use PgBackRest. HA clusters get both from the Patroni configuration template. This role also checks the Patroni configuration store and adds them there if they are missing, because a cluster that was bootstrapped without a repository keeps the/bin/trueplaceholder in the store. It then waits until the leader's Postgres is using the newarchive_command. - Take a full backup only when the repository holds no backup that can restore the cluster running now.
- Create cron entries for scheduled full and differential backups.
Role Dependencies
This role requires the following roles for normal operation:
role_configprovides shared configuration variables to the role.setup_backrestwrites the PgBackRest configuration this role acts through.setup_patroniorsetup_postgresleaves a running cluster for the stanza and the first backup to be taken against.
When to Use
Apply this role at the very end of a deployment, after the cluster is wired together, on both the pgEdge nodes and any dedicated backup servers:
- hosts: pgedge:backup
collections:
- pgedge.platform
roles:
- finalize_backrest
Configuration
This role uses the following parameters from the inventory file:
| Parameter | Use Case |
|---|---|
full_backup_schedule |
Cron schedule for full backups; an empty string installs no entry. |
diff_backup_schedule |
Cron schedule for differential backups; an empty string installs no entry. |
backup_user |
Backup database user (default: backrest). |
backup_password |
Password for the backup database user. |
backup_repo_type |
Decides whether the stanza and the first backup are driven from a pgEdge node or from the backup server. |
backup_repo_user |
OS user the backup server runs PgBackRest as. |
The two schedules are this role's own defaults, because it is the only role that
reads them. Everything else PgBackRest needs — the repository type, path,
encryption and retention — is rendered into pgbackrest.conf by
setup_backrest and read from the file here. See the
Backup Configuration reference for descriptions
and defaults.
How It Works
The role's one decision is whether to take a backup, and it asks the repository
rather than assuming. pgbackrest info reports what the stanza holds; the
stanza is created only when it does not exist, and a full backup is taken only
when the repository holds no backup that can restore the cluster running now.
That matters because of retention. full_backup_count defaults to 1, which
renders repo1-retention-full=1, so a full backup taken against a repository
that already has one expires the previous full backup and the WAL that belongs
to it the moment it completes. A deployment re-run against a cluster whose
repository has been collecting backups for months would otherwise discard the
recovery point it was holding. Because the decision is made from the
repository's contents rather than from where the role sits in a playbook, the
role is safe to apply at any point after the cluster is up.
Two kinds of backup in the stanza do not count. Backups of a cluster the stanza
described before a stanza-upgrade restore that cluster, not this one. And
after a point-in-time recovery, the newest backups can lie on the timeline the
recovery abandoned, past the point where the cluster branched away from it;
PgBackRest restores the newest backup by default, and Postgres cannot follow
the cluster's timeline from there. The role reads the cluster's timeline
history from the zone's primary (a backup server asks the zone's first pgEdge
node, since any member holds the same history), and takes a full backup when the
newest backup ended after the cluster left its timeline. A failover never
causes one, because it branches after every existing backup.
This is what makes the role the step that commits a recovery: see Recovering a Cluster from Backup.
In SSH mode the stanza and the first backup are driven from the dedicated backup
server, which is where the repository lives. In S3 mode there is no server, so
the zone's primary drives them instead. A pgEdge node's pgbackrest.conf names
only its own Postgres, so a replica cannot take a backup: PgBackRest stops with
"unable to find primary cluster". On a non-HA cluster the primary is the zone's
first node. On an HA cluster it is whichever node Patroni reports as leader,
which after a failover is no longer the first node; the role asks
patronictl list before it writes anything, and stops if the zone does not
have exactly one leader among its nodes. The backup database user is created on
the same node, since a replica is read-only.
Repository Identity
Before anything else, the role reads the system identifier from the node's
pg_control and compares it against the identifiers the stanza reports. A
mismatch means this cluster is not the one the repository holds backups for:
archiving would be refused, the scheduled backups would fail nightly, and the
repository's recovery point would stop advancing while continuing to look
healthy. The role stops instead.
An empty repository has no identity to compare, so a genuine first deployment
passes. A cluster rebuilt deliberately during a recovery matches too, because
recover_postgres records it in the stanza's history with
pgbackrest stanza-upgrade first.
setup_postgres asks the same question earlier, before it initializes a data
directory, using the same task file in role_config. The two differ in what an
unreadable repository means: the early check only ever refuses, so it skips a
repository it cannot reach — an SSH repository is not reachable from a pgEdge
node until the backup server has authorized its key. This check gates a decision
to write, so an unreadable repository is fatal here. Concluding "no backups
here" from a repository that could not be read is how a good backup gets
expired.
Readable is judged from the stanza status pgbackrest info reports, not its
exit code, which is 0 even when the repository host is unreachable or its
cipher cannot be decrypted. Only status 0 (ok), 1 (no stanza yet) and 2 (no
backups yet) count. Anything else stops the role, including status 3, the
directories without info files that an interrupted stanza-create leaves
behind.
Artifacts
This role generates and modifies the following files on inventory hosts:
| File | New / Modified | Explanation |
|---|---|---|
{{ backup_repo_path }}/archive/ |
New | WAL archive storage, created by stanza-create (SSH mode; in the bucket for S3). |
{{ backup_repo_path }}/backup/ |
New | Backup storage, created by stanza-create (SSH mode; in the bucket for S3). |
Crontab for backup_repo_user on the backup server (SSH) or postgres on each pgEdge node (S3) |
Modified | Scheduled full and differential backup entries. |
Idempotency
This role is idempotent and safe to re-run. It creates neither a stanza nor a backup that already exists, and cron entries are replaced rather than duplicated. Re-running it against a healthy cluster changes nothing in the repository.