How an update runs

The ordered steps of a guarded install: database backup, image pull by digest, a new app container beside the old one, a health check, then the switchover.

For: Administrators (the admin ability) · Last updated

An install replaces the app and web containers. It never touches postgres or dokku, and it does not change the compose file, volumes, ports or environment. See What an update changes. The design keeps the running deployment serving traffic until the new version has proved itself healthy, and it records every step so the page can show you where things stand.

Phase 1: prepare, while the current version keeps serving#

  1. Backup. The application dumps the Postgres database with pg_dump in custom format to /data/backups/pre-update-<timestamp>.dump, where the timestamp is UTC, for example pre-update-20261003T141500Z.dump. It then reads the dump back with pg_restore --list and gives the file its final name only if it is complete and readable. A failed or timed-out backup stops the install before anything changes. The dump is kept after the install either way.

    /data is the wi-prod-app-data volume. The dump is not shown on the Updates page. See When an update fails for what you can say about it.

  2. Pull. The application pulls the new app and web images by digest, the exact digests the check reported, never by re-resolving the tag. Before it moves any tag, it also tags the image each container is currently running as pre-update-<timestamp>, so the previous images survive an image prune. It then points the running tag (for example latest) at the new images.

  3. Start the new app beside the current one. A new app container, called the "blue" container, starts from the new image next to the current one. It is named like the current container with -blue appended, and it has its own network alias, so the web server keeps sending traffic only to the current app. Database migrations run when blue starts, because the container's entrypoint applies them before it serves.

  4. Health-check blue. The application polls blue's /healthz endpoint every 5 seconds for up to 120 seconds. A container that exits, restarts repeatedly, or never answers fails the install. Because migrations run at start, a migration failure means blue never becomes healthy, and the health check catches it.

  5. Hand off. The application starts a short-lived helper container from the same image and finishes its own work. The helper performs the switchover. This is separate because stopping the old app would kill the very process that was running the install.

Phase 2: the switchover, run by the helper#

The helper performs these steps in order. If one fails, its name appears in parentheses in the failure message on the Updates page.

Step What happens
preflight Checks that everything it needs exists, before it stops anything.
stop_old_app Stops the current app.
remove_old_app Removes it.
rename_blue Renames blue to the canonical app name.
realias_blue Reconnects blue to the network with the app alias.
restart_app Restarts the promoted app, so its database connections are opened on its final network identity.
health_check_app Checks /healthz through the app alias, the same way the web server reaches it.
stop_old_web Stops the current web.
remove_old_web Removes it.
recreate_web Creates web again from the new image.
health_check_web Checks that the new web answers.

web is the only container holding host ports 80 and 443, so the new one cannot start until the old one is gone. That is the brief gap in access you are warned about in the confirmation dialog.

Caddy's certificate volume is carried over to the new web, so the certificate survives updates. Nothing in the process deletes images or the pre-update backup.

Statuses#

The install records one of these statuses as it goes, and the page shows the five working ones as the step list:

backing_up, pulling, starting_blue, health_checking, cutover_in_progress, then success or failed.

Update runs that stop mid-way#

If an install makes no progress for ten minutes and the workflow behind it is no longer alive, for example because the host rebooted partway through, the next status check marks it failed with the message "install did not complete — no active process found, marked failed after timeout". That frees you to start a new install.