When an update fails

What a failed install looks like, what state your deployment is left in, and what is kept for inspection.

For: Administrators (the admin ability) · Last updated

When an install does not complete, the Updates page shows The install did not complete, followed by the failing step in parentheses and the error message, with a Dismiss button:

The install did not complete (health_check_blue): ...

Dismiss it when you have noted the details. You can then check for updates and try again.

What state you are left in#

It depends on how far the install got.

Before the switchover#

These steps run while the current version keeps serving. If any of them fails, your running deployment is untouched:

Failing step Meaning
begin Another install is still active for the company, or the run could not be recorded.
backup The database dump failed, was not readable, or took longer than an hour. Nothing changed.
pull An image could not be pulled by its digest. Check Registry credentials for updates.
start_blue The new app container could not be created.
health_check_blue The new app did not become healthy within 120 seconds, or exited or kept restarting. A common cause is a release that needs a stack change you have not applied: see What an update changes.
launch_cutover The helper container could not be started.

On a failed health_check_blue:

  • The new (-blue) app container is left running, not removed, so you can read its log with sudo docker logs <container name>. The container name is <app container>-blue; find it with sudo docker ps -a.
  • The images that were pulled stay on the box.
  • The current app keeps serving. The new app has already applied its database migrations, so the database may be ahead of the version still serving.

Nothing cleans these up for you. Remove the blue container yourself when you are done inspecting it. A later install replaces a leftover blue container on its own.

During the switchover#

If one of the switchover steps fails (stop_old_app through health_check_web), the containers can be mid-replacement. If replacing web fails after the old one was stopped, the helper tries to put the previous web image back, and the message includes whether that succeeded. Check the state of your containers:

sudo docker compose -p wi-prod ps -a

The helper also records a failure if it receives a stop signal part-way through.

What is kept#

  • The database dump pre-update-<timestamp>.dump in /data/backups inside the app container, which is the wi-prod-app-data volume. It is never deleted automatically.
  • The previous images, tagged pre-update-<timestamp>.
  • The blue container after a failed health check.
  • The helper container that ran the switchover, named wi-prod-update-cutover-<first 8 characters of the run id>. It is not removed, and its log (sudo docker logs <helper name>) holds the detail behind a failed switchover. Find it with sudo docker ps -a.

The Updates page does not display the dump's path. The path follows the pattern above, and the newest file in that directory is the one the latest install made. To list it:

sudo docker compose -p wi-prod exec app ls -l /data/backups

Restoring#

The dump is a standard PostgreSQL custom-format file, verified readable with pg_restore when it was written. The product does not include a restore command. Restoring a dump is a database-administration task that you carry out with PostgreSQL's own tools against the postgres container. Keep copies of the dump you need outside the volume, along with your .env. See Backups and recovery on AWS for the volume-level protection of your deployment.

Other things that look like failures#

  • Check for updates fails immediately. See the error table in Checking for and installing updates.
  • The page shows "Waiting for the app to come back online…" during the switchover. This is normal. It clears when the new app answers.
  • The page shows a red error with a Dismiss button after the connection dropped for a while. Dismiss it and reopen the page. The run's outcome is recorded on the server.