Backups and recovery on AWS

How the CloudFormation deployment protects its root volume, how to check the backups, how to recover a lost instance, and what to arrange on your own server.

For: Administrators operating a deployment · Last updated

Everything that matters lives on one place: the Docker volumes (Postgres, application files, Dokku data, Caddy certificates) and .env, which holds WI_MASTER_KEY. On the AWS path that is the instance's root EBS volume, tagged <NamePrefix>-root-volume.

Protections on the AWS path#

Three independent protections guard the root volume.

  1. Termination protection. The instance has DisableApiTermination and a CloudFormation Retain policy. Deleting the stack fails on the instance until you turn termination protection off yourself. If a stack update replaces the instance, CloudFormation builds the new one first and then fails to terminate the old one, so the old instance and its data stay in place next to a new, empty instance. The Elastic IP and DNS move to the new instance, so the site looks like a fresh install. That is your signal to stop and recover (below), not to start onboarding again.
  2. The root volume outlives its instance. DeleteOnTermination is false. If an instance is terminated on purpose, the volume stays in your account as an unattached, encrypted EBS volume. Delete it yourself once you are sure you do not need it. It continues to be billed until you do.
  3. Daily AWS Backup. A shared plan snapshots every root volume tagged wi-backup=true at 05:00 UTC into the vault wi-backup-vault. Recovery points are kept for BackupRetentionDays (14 by default, set on the primary stack). The root volume is encrypted, and the recovery points keep that encryption.

Check that backups are running in the AWS Backup console under Protected resources, or:

aws backup list-recovery-points-by-backup-vault --backup-vault-name wi-backup-vault

Always read the change set#

Before any stack update, create a change set and read it. If Instance shows Replacement: True, do not execute it unless you mean to rebuild the box. See Install on AWS with CloudFormation.

Recover from a lost instance#

You rebuild the root volume from a backup with the template's SnapshotId parameter.

  1. Find the most recent good recovery point of the lost volume:

    aws backup list-recovery-points-by-backup-vault \
      --backup-vault-name wi-backup-vault \
      --by-resource-type EBS \
      --query 'RecoveryPoints[].[CreationDate,ResourceArn,RecoveryPointArn,Status]' \
      --output table
    

    ResourceArn is the backed-up volume (...:volume/vol-...). For EBS, the recovery point ARN is an EBS snapshot ARN (arn:aws:ec2:<region>::snapshot/snap-...). The snap-... at the end is the snapshot id.

  2. Get a snapshot id to launch from. Either:

    • Use that recovery point's snap-... directly. It is an ordinary EBS snapshot in your account, visible in the EC2 console under Snapshots. Check it first with aws ec2 describe-snapshots --snapshot-ids snap-... and confirm the state is completed.
    • Restore the recovery point to a new volume (AWS Backup console, Protected resources, the volume, Restore, EBS volume, in the instance's Availability Zone), then run aws ec2 create-snapshot --volume-id <restored vol-...> and wait for it to complete. Use this route if the direct one is refused, or if the recovery point is in cold storage (restoring from cold storage can take up to 72 hours).
  3. Launch from it. Update the stack, or create a new one, with SnapshotId=snap-.... EC2 builds the new root volume from that snapshot instead of the AMI's blank one. RootVolumeSizeGiB must be at least the snapshot's size. Keep every other parameter (NamePrefix, BackupVaultName, domain) the same. Updating an existing stack's SnapshotId replaces its instance, which is the point here.

  4. On first boot the startup script runs again on the restored disk. It leaves the existing /opt/wi/.env and its WI_MASTER_KEY alone, refreshes install.sh and find-setup-token.sh, and every container comes back on its own. Check with:

    cd /opt/wi && sudo docker compose -p wi-prod ps
    

    Then load the site and sign in as usual.

  5. Set SnapshotId back to blank only as part of a planned rebuild, because changing it again replaces the instance again. It is fine to leave it set. If the old termination-protected instance or an orphaned volume is still around, delete it only after the recovered box is confirmed good.

Your own server#

The repository provides AWS Backup for the CloudFormation path only. On a server you run yourself, you are responsible for backing up:

  • the Docker volumes wi-prod-data, wi-prod-app-data, and wi-prod-dokku-data (the architecture's convention is one durable volume per host, backed up as a unit);
  • the .env file, kept separately from the volumes, because a restored volume is unreadable without the same WI_MASTER_KEY. See Master key custody.

Arrange this before you rely on the deployment. The wi-prod-caddy-data volume holds your certificates and can be kept as well, but you can reissue certificates if you lose it, subject to the certificate authority's rate limits.

Database backups taken by updates#

Each guarded update takes a database dump before it changes anything. These are not scheduled backups. See How an update runs.